ATS Checker Testing Methodology: Sample Set and Scoring Protocol

Method published 4 August 2026Reviewed 14 August 2026Benchmark status: testing not yet publishedVersion 1.0

This methodology defines how GradVix should test ATS resume checkers using controlled resumes, job descriptions, repeat runs and documented reviewer decisions before publishing any comparative benchmark.

Quick answer: Use synthetic or properly anonymised resumes, preserve one fixed input set, test every tool under equivalent conditions, repeat each run, introduce one controlled change at a time and publish both the findings and the limitations. Do not publish a winner until the evidence can be reproduced.
No benchmark results are claimed on this page. It is the protocol that must be followed before GradVix publishes product rankings, performance scores or “best ATS checker” conclusions.

1. Benchmark objectives

The benchmark should answer narrow, testable questions rather than attempt to prove which product reproduces every employer ATS.

Parsing reliability

Does the tool extract headings, dates, experience, education and skills in a logical order?

Diagnostic quality

Does it distinguish relevance, evidence, formatting, completeness and writing quality?

Explainability

Can a reviewer identify why the result changed and which resume evidence caused it?

Repeatability

Do identical inputs produce materially similar category-level findings?

Controlled sensitivity

Does one genuine correction improve the relevant area without destabilising unrelated areas?

Workflow value

Can a user move from diagnosis to a safer, clearer and more truthful application?

Trust quality

Are privacy, pricing, limits, renewal and refund terms understandable before commitment?

India relevance

Are fresher projects, Indian education, notice periods and local profile workflows interpreted sensibly?

2. Controlled resume sample set

Version 1.0 should use at least twelve synthetic or fully anonymised resume profiles. Each profile must have a written ground-truth record stating what is accurate, deliberately missing, weak, strong or unsupported.

Case Profile Experience level Primary evidence Designed test purpose
R01 Software developer Fresher Final-year Java project, internship Projects, technical skills and fresher completeness
R02 Data analyst Fresher SQL, Excel and Power BI projects Skill evidence versus keyword-only listing
R03 Business analyst 1–2 years Requirements, process mapping, reporting Transferable evidence and role terminology
R04 Application support analyst 3–5 years Incident, SQL and production-support work Operational evidence and role transition
R05 Digital marketer Fresher Campaign project, SEO and analytics Marketing metrics and unsupported claims
R06 Finance analyst 2–4 years Reporting, reconciliation and Excel Domain terminology and quantified outcomes
R07 Customer support professional 3–5 years Ticket handling, SLA and escalation work Impact wording without fabricated numbers
R08 Mechanical engineering fresher Fresher CAD project and industrial training Non-software fresher interpretation
R09 Career switcher to data 5+ years total Operations experience plus new analytics projects Transferable skills and target-role suitability
R10 HR operations professional 2–4 years Onboarding, records and HRIS exposure Process evidence and skill relevance
R11 Sales executive 2–4 years Lead handling, CRM and territory work Metric validation and commercial terminology
R12 Content writer Fresher Portfolio, internship and editorial projects Portfolio evidence and non-technical parsing

3. Resume document variants

Each base profile should be exported into controlled variants. The factual content must remain constant unless the variant is specifically testing an evidence change.

Variant Change Expected observation
V1 Clean single-column, selectable text Baseline parsing and diagnostic result
V2 Two-column layout with the same content Whether reading order or extraction changes
V3 Decorative icons and visual skill bars Whether graphical information is ignored or misread
V4 Image-only or scanned PDF Whether the tool warns about extraction limitations
V5 Non-standard section headings Whether sections are recognised or flagged
V6 Missing project or experience evidence Whether the relevant completeness and evidence areas decline
V7 Keyword-stuffed skills list Whether unsupported terms are rewarded or challenged
V8 One verified correction added Whether only the relevant category improves

4. Job-description controls

Each role should use at least two job descriptions:

  • JD-A: a realistic, suitable vacancy aligned with the candidate’s level.
  • JD-B: a related but materially different vacancy that should create a lower role-match result.

The complete text must be saved with source URL, access date, role title, location, experience range, mandatory conditions and preferred conditions. Remove employer identity only when publication requires anonymisation; do not rewrite the requirements during testing.

Suitability is part of the test. A checker should not hide a major experience or eligibility mismatch merely because the resume contains many related keywords.

5. Standard test procedure

1Freeze inputs

Hash or archive each resume and job-description version.

2Record conditions

Date, browser, account tier, file type and visible product version.

3Run baseline

Submit V1 with JD-A and preserve the complete report.

4Repeat

Run the identical baseline at least three times where permitted.

5Change one factor

Test each document variant or one verified content correction.

6Review and score

Two reviewers compare outputs against the ground-truth record.

6. Benchmark metrics

Parsing accuracyCorrectly extracted required fields ÷ required fields in the ground truth.
RepeatabilitySimilarity of category-level findings across identical runs.
Controlled sensitivityWhether the intended category changes after one verified correction.
Explanation coverageMaterial findings with a location, reason and actionable correction.
Unsupported-suggestion rateRecommendations that would require unverified or false candidate claims.
Workflow completionAbility to revise, preserve versions, recheck and export without data loss.
Metric Calculation or review rule Why it matters
Field extraction Score contact, headings, dates, titles, education, skills and project content against ground truth. Measures basic parsing reliability.
Category stability Compare identical-run findings; flag large unexplained changes. Prevents one-off scores from being treated as dependable.
Issue precision Reviewer confirms whether each flagged issue actually exists. Measures false-positive behaviour.
Missed-issue rate Count known seeded problems not detected by the tool. Measures false-negative behaviour.
Actionability Finding includes affected area, reason, priority and safe correction. Separates diagnostic value from generic advice.
Truth safety Count suggestions that encourage unsupported skills, titles, metrics or outcomes. Protects application credibility.
Trust clarity Checklist for privacy, deletion, pricing, renewal, usage and refund disclosures. Measures commercial and data transparency.

7. Human-review protocol

At least two reviewers should independently assess each report against the ground-truth file. Reviewers should not know the intended brand ranking while coding findings.

  1. Mark each tool finding as correct, partly correct, incorrect or unverifiable.
  2. Mark each known seeded issue as detected, partly detected or missed.
  3. Record whether the recommendation can be followed without inventing information.
  4. Resolve disagreements through a documented adjudication note.
  5. Preserve both the original reviewer decisions and the final resolved decision.

When reviewer judgement remains uncertain, the published result must show the uncertainty rather than forcing a definitive score.

8. Privacy and data controls

  • Use synthetic resumes wherever possible.
  • For real resumes, obtain informed permission and remove names, contact details, identity numbers, addresses and confidential employer information.
  • Do not upload Aadhaar, PAN, passport, bank, medical or authentication information.
  • Record each service’s visible retention, deletion and provider-processing disclosures.
  • Delete benchmark accounts or test data after the retention requirement is complete where the product permits it.

9. Publication and reporting rules

Every published benchmark should include:

  • testing dates and methodology version;
  • tools and account tiers tested;
  • sample size and role coverage;
  • document and job-description variants;
  • raw category results or a sufficiently detailed appendix;
  • reviewer method and disagreement handling;
  • known limitations, failures and unavailable features;
  • commercial relationships or affiliate interests;
  • separate product-quality scores and resume ATS scores; and
  • a correction log when a product or result changes.
Do not publish “best ATS checker” from one resume and one job description. A small demonstration can illustrate a workflow, but it cannot support a broad market ranking.

10. Required limitations

  • The benchmark cannot access private employer configurations.
  • AI-backed tools may change without a visible version number.
  • Scores may vary by account tier, file type, traffic conditions or product update.
  • Reviewer judgement can introduce subjectivity despite a ground-truth record.
  • A test set cannot represent every Indian role, language, industry or experience level.
  • Strong benchmark performance does not guarantee an interview or hiring result.

Frequently asked questions

Has GradVix already published benchmark results?
No. This page defines the protocol. Results should be published only after the controlled tests, repeat runs and human review are complete.
Why use synthetic resumes?
They provide known ground truth, allow deliberate test defects and reduce privacy risk. Properly anonymised real resumes can supplement them with consent.
Why repeat identical runs?
Some AI-assisted systems can vary. Repeat runs show whether category-level conclusions remain stable enough to be useful.
Can the benchmark prove which tool matches employer ATS software?
No. It can compare public checker behaviour, parsing, explanations and workflow. It cannot reproduce every employer’s private configuration.
How often should results be refreshed?
Refresh after material product changes and at a clearly stated review interval. Preserve the previous dated results rather than silently replacing them.
Scroll to Top