This methodology defines how GradVix should test ATS resume checkers using controlled resumes, job descriptions, repeat runs and documented reviewer decisions before publishing any comparative benchmark.
1. Benchmark objectives
The benchmark should answer narrow, testable questions rather than attempt to prove which product reproduces every employer ATS.
Does the tool extract headings, dates, experience, education and skills in a logical order?
Does it distinguish relevance, evidence, formatting, completeness and writing quality?
Can a reviewer identify why the result changed and which resume evidence caused it?
Do identical inputs produce materially similar category-level findings?
Does one genuine correction improve the relevant area without destabilising unrelated areas?
Can a user move from diagnosis to a safer, clearer and more truthful application?
Are privacy, pricing, limits, renewal and refund terms understandable before commitment?
Are fresher projects, Indian education, notice periods and local profile workflows interpreted sensibly?
2. Controlled resume sample set
Version 1.0 should use at least twelve synthetic or fully anonymised resume profiles. Each profile must have a written ground-truth record stating what is accurate, deliberately missing, weak, strong or unsupported.
| Case | Profile | Experience level | Primary evidence | Designed test purpose |
|---|---|---|---|---|
| R01 | Software developer | Fresher | Final-year Java project, internship | Projects, technical skills and fresher completeness |
| R02 | Data analyst | Fresher | SQL, Excel and Power BI projects | Skill evidence versus keyword-only listing |
| R03 | Business analyst | 1–2 years | Requirements, process mapping, reporting | Transferable evidence and role terminology |
| R04 | Application support analyst | 3–5 years | Incident, SQL and production-support work | Operational evidence and role transition |
| R05 | Digital marketer | Fresher | Campaign project, SEO and analytics | Marketing metrics and unsupported claims |
| R06 | Finance analyst | 2–4 years | Reporting, reconciliation and Excel | Domain terminology and quantified outcomes |
| R07 | Customer support professional | 3–5 years | Ticket handling, SLA and escalation work | Impact wording without fabricated numbers |
| R08 | Mechanical engineering fresher | Fresher | CAD project and industrial training | Non-software fresher interpretation |
| R09 | Career switcher to data | 5+ years total | Operations experience plus new analytics projects | Transferable skills and target-role suitability |
| R10 | HR operations professional | 2–4 years | Onboarding, records and HRIS exposure | Process evidence and skill relevance |
| R11 | Sales executive | 2–4 years | Lead handling, CRM and territory work | Metric validation and commercial terminology |
| R12 | Content writer | Fresher | Portfolio, internship and editorial projects | Portfolio evidence and non-technical parsing |
3. Resume document variants
Each base profile should be exported into controlled variants. The factual content must remain constant unless the variant is specifically testing an evidence change.
| Variant | Change | Expected observation |
|---|---|---|
| V1 | Clean single-column, selectable text | Baseline parsing and diagnostic result |
| V2 | Two-column layout with the same content | Whether reading order or extraction changes |
| V3 | Decorative icons and visual skill bars | Whether graphical information is ignored or misread |
| V4 | Image-only or scanned PDF | Whether the tool warns about extraction limitations |
| V5 | Non-standard section headings | Whether sections are recognised or flagged |
| V6 | Missing project or experience evidence | Whether the relevant completeness and evidence areas decline |
| V7 | Keyword-stuffed skills list | Whether unsupported terms are rewarded or challenged |
| V8 | One verified correction added | Whether only the relevant category improves |
4. Job-description controls
Each role should use at least two job descriptions:
- JD-A: a realistic, suitable vacancy aligned with the candidate’s level.
- JD-B: a related but materially different vacancy that should create a lower role-match result.
The complete text must be saved with source URL, access date, role title, location, experience range, mandatory conditions and preferred conditions. Remove employer identity only when publication requires anonymisation; do not rewrite the requirements during testing.
5. Standard test procedure
Hash or archive each resume and job-description version.
Date, browser, account tier, file type and visible product version.
Submit V1 with JD-A and preserve the complete report.
Run the identical baseline at least three times where permitted.
Test each document variant or one verified content correction.
Two reviewers compare outputs against the ground-truth record.
6. Benchmark metrics
| Metric | Calculation or review rule | Why it matters |
|---|---|---|
| Field extraction | Score contact, headings, dates, titles, education, skills and project content against ground truth. | Measures basic parsing reliability. |
| Category stability | Compare identical-run findings; flag large unexplained changes. | Prevents one-off scores from being treated as dependable. |
| Issue precision | Reviewer confirms whether each flagged issue actually exists. | Measures false-positive behaviour. |
| Missed-issue rate | Count known seeded problems not detected by the tool. | Measures false-negative behaviour. |
| Actionability | Finding includes affected area, reason, priority and safe correction. | Separates diagnostic value from generic advice. |
| Truth safety | Count suggestions that encourage unsupported skills, titles, metrics or outcomes. | Protects application credibility. |
| Trust clarity | Checklist for privacy, deletion, pricing, renewal, usage and refund disclosures. | Measures commercial and data transparency. |
7. Human-review protocol
At least two reviewers should independently assess each report against the ground-truth file. Reviewers should not know the intended brand ranking while coding findings.
- Mark each tool finding as correct, partly correct, incorrect or unverifiable.
- Mark each known seeded issue as detected, partly detected or missed.
- Record whether the recommendation can be followed without inventing information.
- Resolve disagreements through a documented adjudication note.
- Preserve both the original reviewer decisions and the final resolved decision.
When reviewer judgement remains uncertain, the published result must show the uncertainty rather than forcing a definitive score.
8. Privacy and data controls
- Use synthetic resumes wherever possible.
- For real resumes, obtain informed permission and remove names, contact details, identity numbers, addresses and confidential employer information.
- Do not upload Aadhaar, PAN, passport, bank, medical or authentication information.
- Record each service’s visible retention, deletion and provider-processing disclosures.
- Delete benchmark accounts or test data after the retention requirement is complete where the product permits it.
9. Publication and reporting rules
Every published benchmark should include:
- testing dates and methodology version;
- tools and account tiers tested;
- sample size and role coverage;
- document and job-description variants;
- raw category results or a sufficiently detailed appendix;
- reviewer method and disagreement handling;
- known limitations, failures and unavailable features;
- commercial relationships or affiliate interests;
- separate product-quality scores and resume ATS scores; and
- a correction log when a product or result changes.
10. Required limitations
- The benchmark cannot access private employer configurations.
- AI-backed tools may change without a visible version number.
- Scores may vary by account tier, file type, traffic conditions or product update.
- Reviewer judgement can introduce subjectivity despite a ground-truth record.
- A test set cannot represent every Indian role, language, industry or experience level.
- Strong benchmark performance does not guarantee an interview or hiring result.