FieldFinch – Project Learning Blog

9

October

2026

4/5 (1)

We created FieldFinch, an AI driven matchmaking platform that helps early-career researchers find the right expert beyond their own network: you describe the knowledge which your research requires and the platform provides experts from different fields with real publications as proof for every match.

My main contributions to the project were the problem & opportunity, the evaluation (simulation slide) and the second slide of the final recommendation in which we used the project findings to make decisions about what parts of the prototype to implement, change, test, scale or stop. These parts where closely connected. The problem & opportunity provided the initial explanation for why an AI driven academic matching platform could provide value, the simulation then tested if AI is capable of improving this matching process and the recommendation turned these results into practical next steps. Furthermore, I also contributed a lot to the overall consistency of the slideshow. I rewrote several parts to clarify certain parts of the project and adjusted the formatting and layouts so the slideshow looked like one coherent project.

The most difficult part of the project for me was deciding how we could meaningfully evaluate our concept, just claiming that AI would create better matches would not have provided convincing evidence. FieldFinch already had a prototype for the user interface but this did not prove if the AI matching functionality could actually perform the task that we had in mind. We also could not simply observe if researchers were finding useful experts through our platform without an existing user base. I therefore decided to test our claim with a small simulation which I had Claude run on OpenAlex data. The simulation used 5 academic collaborations from 2023 which were all written by two authors, each from a different field of research, who had also not worked together before. The simulation then picked one author as the “seeker” and solely searched for work published before 2023. Next, I checked if each search method found experts who were on-topic, if they were in the collaborator’s field and lastly if the real collaborator was found. The most important part of providing the evaluation was finding a meaningful baseline. Only using a simple keyword search in the “seeker’s” own words found zero on-topic experts, which would make our AI matching seem much more powerful. An informed keyword search using the terminology of an expert, however,  reached 68% on-topic against 84% on-topic for FieldFinch v1. After comparing FieldFinch to a fairer baseline we changed our claim, I had evidence that the AI was mostly useful for providing more reach, not relevance. This led me to the final conclusion that FieldFinch could find five times more experts from other fields than is currently possible through searching by keywords.

I learned most from the failures that the simulation produced. Claude built a second version of the simulation which split each research question into separate searches for the method and the field of research. Here I identified progress in the reach for cross-field research, which rose from 20% to 44%. Though the percentage of on-topic experts fell from 84% to 0%. For example, for one of the questions which was about patient no-shows, v2 suggested a paper on security in Java software. This is why I decided to keep v1, because having more reach without finding relevant research is quite useless for a researcher. Another failure was that no version of FieldFinch could find the real co-author of the paper in any of the five cases. The researchers earlier work was focused on general research methods, not the seeker’s specific problem which is why it never reached the top results. The data that we received also complicated our own problem statement. Several pairs of researchers we first looked at had already published together. This supports our claim that collaborations in scientific research are driven by existing networks and conversely shows that is very difficult for any matching tool to replace them. There were also some limitations to the simulation. Firstly, five cases is a very small sample size. Secondly, I let Claude run the entire simulation, including writing the search query. It wrote those queries with the real papers in mind which may have inflated the results of FieldFinch v1. That is why we eventually decided run a pilot with go/no-go targets instead of a full launch of the application.

Being responsible for both the evaluation and part of the recommendation made me see how evidence from the simulation changed our decisions. When writing the recommendation, I purposefully connected all decisions to a result from the simulation, so that the next steps would be based on real results instead of an opinion. Furthermore, I decided to implement FieldFinch v1 because it found a lot more cross-field experts at similar relevance. For the pilot launch, I made the choice that the tool should provide a source link for every claim that it makes to increase its trustworthiness and decided that the tool should match on whole expert profiles with embeddings, since it missed real co-authors in both v1 and v2. The evidence provided by the simulation also changed my assumption that AI would make finding experts much simpler. From v2, I could conclude that AI widens the search beyond a researcher’s own field, but it still misses expertise that can be transferred. I chose not to hide the weak result of the tool not matching 1 of the 5 cases with the co-author. Instead I used this as a reason to pilot FieldFinch at one research office first, with certain go/no-go targets.

The project taught me that during collaboration, dividing the work is the easy part. The hard part is making all those parts read as one project and making all those parts connect to be as convincing as possible. By rewriting sentences and changing the formatting I learned that the integration of individual work needs a clear owner, otherwise a project will end up looking uncoherent. If I would do the project again, I would assign the final round of proofreading and correcting as a separate task.

Starting out, I believed AI was a product on its own: you can give it an instruction and it matches people on its own. The project showed me, however, that there is much more to connecting the right people than just “letting AI match”. The instructions you give the AI tool decides a lot about the outcome, a small change in how v2 formulated the request changed relevant experts to irrelevant ones. It also made me aware that there are areas with clear limits to what AI can do, like matching junior researchers who do not have any publications yet. On the other hand, I was also very impressed by what AI could do, it ran our simulation and made a working prototype without writing any code ourselves. This made me realize that if we could make a protype this quickly, that competitors such as ResearchGate or Elsevier can create a similar matching tool with their existing userbase and data as well.

Please rate this

Leave a Reply

Your email address will not be published. Required fields are marked *