Job-related evidence with a human path

Evaluate AI candidate screening from job analysis through candidate review.

AI screening evaluation starts with the job, not the model. Define the construct and decision, use structured and inspectable administration, examine validity and error evidence, test accessibility and accommodations, and preserve meaningful human review and contestability.

Last substantive review: August 7, 2026 · General evaluation guidance, not legal advice.

Locate consequence

Identify what the screening output changes in the hiring process.

Screening can mean eligibility questions, resume ranking, work samples, structured interviews, conversational assessment, transcription, summarization, or a recommendation. Write whether the output informs review, orders a queue, advances a person, rejects them, or triggers another assessment. The same model output creates different consequences under different employer use.

Map inputs, generated or inferred features, score or narrative output, thresholds, human reviewers, and downstream status changes. Ask whether a person can be rejected solely from the system, whether a reviewer sees underlying evidence, and whether the candidate can request another path. A human who only confirms a score under time pressure may not provide meaningful oversight.

Determine which jurisdictions, roles, and candidate populations are in scope before applying a legal or professional framework. The NYC AEDT materials and EU AI Act explanations illustrate how definitions and intended purpose matter; this guide does not determine coverage. Record the verified source date and seek qualified advice for the actual deployment.

  • Input. Resume, application answer, assessment response, audio, video, behavioral data, or inferred feature.
  • Construct. The job-related knowledge, skill, ability, competency, or eligibility condition intended to be measured.
  • Output. Score, label, rank, summary, recommendation, explanation, or automatic status change.
  • Consequence. Review priority, additional test, advance, rejection, communication, or final selection support.

Start from work

Connect every scored dimension to current job evidence.

Use a current job analysis to identify important tasks, competencies, context, and minimum qualifications. The OPM job-analysis guidance describes this foundation, while EEOC materials address job-related selection procedures. A generic model score such as "leadership," "culture fit," or "communication" is not self-validating; define the behavior, why it matters for this role, and what evidence can support it.

Distinguish minimum eligibility, trainable skill, preference, and speculative predictor. Avoid treating education, employer prestige, career continuity, accent, facial behavior, typing style, or vocabulary as a competency without a defensible link to the work. Proxy variables can appear objective while measuring access or background rather than ability to perform the job.

Freeze the role version and rubric before evaluating candidates. Record who approved each dimension, how it is weighted or gated, acceptable alternative evidence, and how missing information is handled. Revalidate when duties, level, location, technology, or decision use changes.

  1. 01
    Analyze the job. Identify important tasks, competencies, context, and consequences using current role information.
  2. 02
    Define the construct. Describe the capability in observable terms and separate it from convenient proxies.
  3. 03
    Choose evidence. Specify which answers, work samples, or experiences can support the construct and acceptable alternatives.
  4. 04
    Approve the rubric. Set questions, anchors, thresholds, unknown handling, and reviewer authority before live use.

Comparable opportunity

Structure the core questions and scoring without erasing necessary accommodation.

Structured interviews use predetermined job-related questions and common rating standards so candidates receive comparable opportunities to provide evidence. An AI conversation can still be structured: define core questions, permitted probes, time behavior, response channels, scoring anchors, and the conditions under which a human follows up.

Consistency is not identical wording at any cost. A reasonable accommodation, clarification, language support, or recovery from a technical failure may require a different path. Record the change and preserve the construct being measured rather than penalizing the candidate for the delivery mechanism. ADA.gov warns against tests that measure disability instead of job skill.

Test prompt and conversation variation. Re-run semantically equivalent responses, order changes, irrelevant details, concise and verbose answers, uncertainty, non-native language patterns, interruptions, and adversarial or nonsensical input. Inspect whether scoring remains anchored to role evidence or drifts toward style, confidence, or demographic proxy.

Structure to define before an AI-assisted screen
ElementDefineTest
QuestionJob-related purpose, required wording, and allowed clarification.Whether variants change the construct or candidate opportunity.
ProbeWhen the system may ask for detail and when it must stop.Over-questioning, leading prompts, and unequal depth.
RatingBehavioral anchors, evidence rules, unknown state, and threshold.Agreement, borderline cases, and unsupported inference.
DeliveryTime, channel, language, accessibility, interruption, and recovery.Whether interface behavior affects the score.
ReviewWhat evidence the human sees and what they can correct or override.Automation bias, time pressure, and auditability.

Evidence for interpretation

Ask whether the evidence supports the score's intended use.

Validity concerns the interpretation and use of the screening output for the stated decision. Request the argument connecting job analysis, construct, content, response, scoring, threshold, and relevant outcome. Evidence for one occupation, language, population, or use may not transfer to another. A general LLM benchmark does not validate an employment screen.

Reliability and consistency support but do not replace validity. Test repeated scoring, reviewer agreement, model-version movement, and sensitivity to irrelevant changes. A perfectly consistent measure of the wrong construct remains unsuitable. Conversely, legitimate open-ended evidence can contain uncertainty that should be represented rather than hidden behind excessive decimal precision.

Review the sample and label source. Who decided which answers were strong? Were raters trained and blinded? Were disagreements adjudicated? Are protected or relevant subgroups represented well enough for the reported analysis? Are thresholds chosen before or after seeing outcomes? Request limitations and negative results alongside the aggregate accuracy number.

  • Content evidence. Does the assessment represent important parts of the work and competency definition?
  • Response process. Do candidates and the system engage with the question as intended?
  • Internal consistency. Are ratings and items coherent without collapsing distinct competencies?
  • Relations to outcomes. Do scores relate to relevant external evidence under an appropriate design?
  • Consequences. What errors, exclusions, burdens, or adaptations arise from the chosen use?

Look beyond average accuracy

Examine selection impact and error patterns without turning one ratio into a compliance verdict.

Review advancement, score, error, and missing-data patterns for relevant groups where lawful, appropriate, and statistically supportable. The Uniform Guidelines and official NYC materials illustrate that impact analysis has defined contexts and methods. A vendor's generic fairness dashboard or one favorable ratio cannot determine compliance for the buyer's use.

Inspect false positives and false negatives alongside selection rates. A system can produce similar aggregate rates while making different kinds of errors or measuring a proxy differently. Examine intersectional and accessibility-sensitive cases where sample design permits, and record when data is unavailable or too sparse for a stable estimate.

Investigate causes across job definition, question content, training data, labels, missingness, interface, transcription, language, scoring model, threshold, reviewer behavior, and downstream use. Mitigation should address the failure rather than tune a number until one report passes. Re-test after mitigation and monitor in operation.

Candidate access

Test the complete process for accessibility, accommodation, and alternative paths.

Evaluate the candidate journey from notice through completion, correction, and human contact. Test keyboard use, focus order, labels, status messages, text alternatives, contrast, reflow, time limits, captions or transcripts, screen-reader behavior, error recovery, mobile use, and low bandwidth. WCAG 2.2 provides testable web criteria but does not cover every hiring-process need.

Explain the technology and evaluated information early enough for a candidate to decide whether to request an accommodation. Make the request path easy to find, confidential as appropriate, and operationally staffed. Test that requesting or receiving an accommodation does not itself reduce the score or reveal unnecessary disability-related information to evaluators.

Provide an alternative that measures the same job-related construct when the standard path creates a barrier. Do not simply remove the candidate from consideration or substitute an unrelated test. Technical support is not the same as accommodation ownership, and a chatbot that routes in circles is not a human escalation path.

  1. 01
    Inform. Describe the technology, task, timing, evaluated information, and available support before the screen.
  2. 02
    Request. Offer a clear accommodation route that reaches an accountable person.
  3. 03
    Adapt. Use an accessible or alternative method that preserves the job-related construct.
  4. 04
    Verify. Test the complete adapted journey, scoring, records, and downstream review rather than only the interface.

Meaningful control

Give reviewers evidence, time, authority, and a reason to disagree.

Human review is meaningful when the reviewer understands the system's role, can inspect candidate evidence, knows limitations, has time to evaluate, can change the result, and is accountable for the next action. A mandatory confirmation click or unexplained score does not meet that practical standard.

Design the interface to separate source evidence, model inference, rubric rating, confidence or uncertainty, and final human decision. Ask reviewers to give a reason for material overrides and sample non-overridden cases for automation bias. Monitor whether reviewers increasingly accept recommendations without reading evidence as volume grows.

Give candidates and internal users routes to contest, correct, or escalate consequential errors. Preserve the original output and the corrected record so the organization can learn without leaving harmful data active. Communicate outcomes and limits in language appropriate to the process; do not expose proprietary internals when a clear job-related explanation is possible.

  • Evidence access. The reviewer can trace a rating or recommendation to the candidate response and rubric.
  • Decision authority. The reviewer can correct, override, pause, or request another method.
  • Time and training. The workflow does not turn review into rubber-stamping under an unrealistic queue.
  • Contestability. Candidate and user concerns reach an owner and can change relevant records or actions.

Live operation

Monitor score movement, candidate impact, overrides, and process change together.

After deployment, monitor score distributions, advancement, error samples, reviewer disagreement, overrides, missing data, accommodation use, technical failures, candidate feedback, completion, and downstream interview evidence. Segment by role, version, language, channel, and other relevant dimensions where analysis is lawful and supported.

Version prompts, questions, rubrics, thresholds, models, transcription, interface, and integrations. A vendor update can change candidate opportunity even when the product name is unchanged. Establish change notification and re-testing triggers in the agreement and operating process.

Create incident paths for inaccessible sessions, incorrect status changes, corrupted or exposed data, repeated questions, unsupported inferences, group performance concerns, and candidate complaints. Preserve evidence, contain impact, correct records, communicate with affected people where appropriate, and decide whether the system may resume under narrower controls.

Screening monitoring signals and investigation questions
SignalInvestigatePossible action
Score distribution shiftRole mix, model, rubric, prompts, language, data, and candidate population.Pause threshold use, sample cases, or revalidate.
Override changeReviewer training, queue pressure, system quality, and automation bias.Review cases, interface, staffing, and rubric.
Completion or accommodation issueInterface, notice, timing, support, alternative path, and downstream scoring.Fix access, offer reassessment, and limit exposure.
Group or error concernSample, labels, missing data, threshold, construct, and workflow use.Escalate qualified review, contain use, and retest.
Integration incidentIncorrect writes, retries, duplicates, permissions, and audit trail.Disable writes, reconcile records, and use fallback.

Evidence register

Sources used, and what they cannot prove.

professional practice

U.S. Office of Personnel Management: Job Analysis

A practical account of job analysis as the foundation for defining tasks, competencies, and assessment content.

Evidence limit: Federal personnel practice does not by itself validate a private-sector role brief or every automated assessment.
professional practice

U.S. Office of Personnel Management: Structured Interviews

Guidance on using predetermined job-related questions, consistent administration, and common rating standards.

Evidence limit: Structure improves comparability but does not guarantee validity, fairness, accessibility, or a correct hiring decision.
technical standard

World Wide Web Consortium: Web Content Accessibility Guidelines 2.2

Testable web-content accessibility criteria across perceivability, operability, understandability, and robustness.

Evidence limit: WCAG conformance covers web content and does not by itself prove that an end-to-end hiring process is accessible or lawful.
government guidance

European Commission: Navigating the AI Act

Current Commission explanations of scope, risk classification, employment use cases, obligations, and implementation timing.

Evidence limit: The FAQ is explanatory, timing can change, and classification depends on intended purpose and the facts of a deployment.
first-party research

When AI Meets Recruiting: Opportunities, Challenges, and Future Directions

The Metix literature review maps sourcing, matching, and assessment across a recruitment lifecycle and identifies bias, explainability, feedback, and oversight as open deployment questions.

Evidence limit: This first-party review does not validate a screening construct, product, employer process, or legal outcome.

Relationship disclosure: OpenJobs AI is now Metix AI. Metix material is labeled first-party and is not treated as independent validation.

Questions teams ask

Frequently asked questions

Is resume ranking an employment selection procedure?

Its treatment depends on the facts, jurisdiction, definitions, and how the employer uses it. Record the actual decision effect and obtain qualified advice rather than relying on a product label.

Does consistent AI scoring prove the screen is valid?

No. Consistency can support reliability, but validity requires evidence that the score interpretation and use relate appropriately to the job and decision.

Is WCAG conformance enough for an accessible screening process?

No. WCAG addresses web content. The full process also needs timely notice, accommodation handling, suitable alternatives, human support, and downstream treatment.

Can a human reviewer fix every AI screening risk?

No. Review can help only when the person has evidence, competence, time, authority, and a functioning correction path. Some unsuitable constructs or inaccessible methods should not be used.