How we tested Ema Recruiter against the tool your team already pays for - and what it returned

Varied requisitions. One rubric, frozen before a single candidate was scored.
Every recruiting product claims better search. Almost none of them show their work.
So we ran the test. We took real requisitions and put each one - unchanged - through Ema Recruiter and through the incumbent: the seat-based platform most enterprise talent teams already have, already pay for, and already work inside every day. Every candidate in each tool's top 15 was scored, 1,350 in all, against a rubric written before any scoring began.
We chose that comparison because it's the only one a buyer actually faces. If a new tool can't beat what's already on the recruiter's second monitor, it isn't worth a login. Here is the method, the scoring, and the result.
Objective: Identify which system puts more candidates in front of a recruiter that they would actually send to the hiring manager. Everything below is in service of that one question.
What we ran and the methodology
A requisition set built deliberately to look like what recruiters actually type, spread across various domains and kinds of criteria. Two-word searches, boolean strings, conversational briefs ("we need someone who's owned a $10M performance budget"), and full 480-word job descriptions. Roles spanned backend engineering, private equity, ICU nursing, reservoir engineering, hotel general management, HVAC field service, litigation and more - across the US, Seoul, São Paulo, London and the Caribbean.

Top 15 displayed candidates from each tool, in the order each tool displayed them. For Ema, that's the default view a recruiter lands on: the shortlist ranked by scorecard score, with search filters verified empty on every run. Nothing hand-tuned.
Scored by a frontier reasoning model, against a rubric frozen before scoring started. Every deduction is recorded against a named candidate, so any individual call can be re-examined.
Only fields both tools display were scored. If one showed something the other didn't, it came out of the rubric entirely.
The two questions
Sourcing quality isn't one question, and treating it as one is how benchmarks get gamed.
Question 1 - Did it deliver what the requisition asked for? Every requirement quoted verbatim from the req, marked required or preferred, and each candidate scored Yes / Maybe / No against every one. Literal, checkable, unforgiving.
Question 2 - Would a recruiter actually put this person in front of the hiring manager? One holistic judgment per candidate, including everything the requisition didn't say. A candidate can tick every stated box and still fail this, on five counts:
- Is the function right? A telephone-triage nurse against an ICU brief. A product manager against an engineering brief. An investor-relations VP against a private-equity investing brief.
- Is the level right? A partner against an associate brief. An SDR against an account-executive brief. An assistant controller against a controller brief.
- Is the context right? In-house counsel against a private-practice brief. Agency-side against an in-house brief. Residential against commercial.
- Are they a real, current professional? Not a student, an intern, a summer associate, a retiree, or an unpaid club role presented as a job.
- Is the record usable? A real employer name, not a blank field — and not the same person listed twice.
The second question is the one recruiters are actually asking when they scan a list. The first is the one every keyword engine optimizes for.

Ema is ahead on both questions, on 80% of requisitions where both tools returned results, and by thirteen points overall.
The second row is the one a working recruiter feels. Read it as candidates rather than percentages: Ema's 91% means about 13 of every 15 profiles are worth a look, where the incumbent's 76.8% means about 11. Two more dead profiles in every shortlist a recruiter opens - on every search, every req, every week. That's the difference between a list you work and a list you triage.

Three patterns showed up repeatedly in what the incumbent returned:
- Level drift. Against a director-level SaaS sales brief it returned eleven individual contributors. Against an enterprise account-executive brief in Austin, an SDR and an account associate. Against a controller brief, assistant controllers.
- Function drift. Against a payments-engineering brief, four product managers. Against a head-of-performance-marketing brief written conversationally, it degraded into consultants and generalist specialists.
- Empty results. On an enterprise sales search in Seoul, it converted a stated preference into two hard filters and returned no candidates at all. Ema returned fifteen, all in Seoul, all at enterprise software firms.
By contrast, on that same set: all fifteen of Ema's corporate-lawyer results were corporate lawyers at the firms specified. All fifteen private-equity VPs were genuine private-equity investors, not adjacent VP titles. All fifteen supply-chain results sat at the three named employers.
Honest callout: Candidate profiles and each tool’s reasoning taken at face value
- Every candidate card was taken at face value. No profile was checked against reality - a stale title scores exactly the same as a current one, and a person who left that employer last year and didn’t update their profile still counts as being there.
- Each tool's own claims were taken at face value too. Scores, relevance labels and criteria ticks were read as displayed, not audited. Both tools were extended the same courtesy, so the comparison holds - but the absolute numbers describe what the tools say they found.
What produces a shortlist like that?
None of this is the search box working harder. Four things about how Ema Recruiter is built show up directly in the numbers above.
- One search across 1B+ profiles and 30+ sources. Ema pulls from everywhere you work - your ATS, your CRM, licensed data providers and the open web - in a single query. Coverage is why a search in Seoul or São Paulo returns fifteen genuine local candidates instead of an empty page.
- Natural language search, no boolean required. Every requisition in this benchmark went in exactly as written — two-word searches, a conversational brief, a 480-word job description - with no translation into search syntax. Nothing was lost turning a hiring manager's sentence into a query string.
- Every candidate scored and ranked against your role criteria, with the rationale attached. Ema evaluates candidates automatically on a 0-5 scale and reorders the list as scores generate. The top 15 a recruiter sees are the best fifteen of thousands, not the first fifteen the index returned - which is the single biggest reason 91% of them are presentable.
- Bot résumés, weak profiles and duplicates are filtered out before they reach the shortlist. Two of the five ways a candidate failed our second question - a blank or placeholder employer, and the same person listed twice - are conditions Ema screens for automatically. They cost the incumbent points across the study.
Underneath all of it, scoring runs on evidence rather than identity: PII is redacted before scoring and demographic signals are stripped from ranking, with a complete audit trail behind every decision.
Run your hardest requisition against both
Take the req that has been open longest, the one your team keeps re-running, and put it through Ema Recruiter and your current stack side by side. Then count how many of the first fifteen you would actually send to the hiring manager. That number is the only benchmark that matters to you.
