Lab · September 29, 2026
How we check that the right people come first
Seven musicians in our test pool sing to their own guitar and produce the tracks themselves. All seven can do the job.
The person asking writes short messages and wants polished work on the first pass. By that measure, the closest of the seven scores almost twice as high as the furthest.
Folkso has to put that one near the top. We check it on 1,000 searches where the right answer is known in advance, with the same code that answers you in ChatGPT and Claude.
What you get
98%
The first person you see can do what you asked.
96%
One of your three best matches is already in the top three.
93%
All five people in the results can do the job.
99%
If someone who fits exists, Folkso shows them.
A best match fits both the task and the way you like to work. The numbers are from 27 September 2026, on 300 searches that took no part in tuning.
How we got here
First person can do the job
- Random order
- 21.4%
- Before 27 September
- 78.5%
- Now
- 98.1%
Best match in the top three
- Random order
- 11.3%
- Before 27 September
- 57.4%
- Now
- 76.4%
Best match comes first
- Random order
- 4.6%
- Before 27 September
- 28.5%
- Now
- 48.9%
Someone who fits gets shown
- Before 27 September
- 86.7%
- Now
- 99.0%
A change counts as better only if it beats the old version by more than 5 points, the margin on 300 searches.
Why you can trust these numbers
- The test pool has 100 fictional people and 250 fictional people searching. Each one starts as a hidden card with a job, a city, a character, values and tastes.
- Every niche has twins. Five or more people with the same skills and different characters, so search has to look past the job title.
- Language models wrote the profiles and the 1,000 searches, the same kind of searches ChatGPT and Claude send to Folkso.
- The right answer comes from the hidden cards and counts only what a reader can see in the profile.
- 700 searches tune the weights. The other 300 are held out, and every number on this page comes from them.
- Every run is saved with its date, code and results.
How your search runs
- Filters. Only people open to what you need, with a budget in your currency and unit, in your city when it matters. City names work in any language.
- Meaning. Folkso compares your task with every profile by meaning and by words, and keeps the 50 closest.
- Reading. A model reads those 50 profiles in full, without contacts. Can the person do the job? Are they looking for the same role themselves? How close are they to you? What do you share?
- Second look. The first six are read again, one at a time. Only those who pass reach you.
- Nobody fits. Folkso relaxes the soft conditions and says what changed. If there is still nobody, it tells you so.
Your contacts and theirs never reach the model.
Also measured
- No founders in your way. A founder who needs a React developer doesn't show up when you search for one. With the model it happened 0 times in 565 searches.
- Any language. A profile written in another language is found almost as well. Best match first is 47.2% against 50.3% in the same language.
- Speed. Reading and the second look add 0.7 seconds at the median.
- Noise. Two identical runs of the model disagree a little, so we don't count differences smaller than that.
What didn't work
We publish the misses too. Each of these looked promising and lost on the held-out searches.
| Idea | What happened |
|---|---|
| Reading all 50 profiles in one request | 18% of searches failed and ran without the model |
| Character traits as fixed labels | +3 points, within the margin |
| Describing character to the model in words | best match first fell from 49% to 45% |
| Tuning the weights on a grid | +8 points while tuning, +2 on held-out searches |
| Letting your AI pick from ten people | +1.5 to +2.3 points |
Where the test stops
The people here are fictional. Real profiles are shorter and messier, with typos and two languages in one text.
Whether two people work well together shows up in their first conversation. That is why the final choice stays with you.
Next come live outcomes. Accepted requests and “not the one” with a reason will tune the order further.
Folkso is free. Connect it once, and the next time a task needs a person, your AI finds them right in the conversation.