A 4-Prompt AI Resume Screening System Caught Every Planted Trap in 96 Seconds
Resume screening is a high-volume, low-consistency judgment task that burns two full days per hundred applicants. A prompt chain that enforces citation-backed scoring and a mandatory Pending tier shifts the AI's role from black-box decider to auditable triage tool, cutting screening time to minutes while leaving the final call with a human.
Manual resume screening drifts in standards, misses timeline gaps, and leaves no audit trail. A four-step system built in TraeWork replaces that with a fixed scorecard, a batch-screening pass that demands original-text citations for every judgment, and a triage into Recommend, Pending, and Do Not Recommend tiers. The workflow was validated against six simulated resumes with deliberate traps—overlapping dates, skill inflation, and unsupported claims—and caught every one. Batch screening six resumes took 1 minute 36 seconds, and the same session produced eight customized interview questions for the top candidate plus phone-review scripts for the Pending tier. The full set of four prompt templates is published for direct reuse, with the core rule being that speculation is forbidden and every conclusion must cite the source text.
The system's real value is not speed but auditability: a scorecard plus mandatory citations turns an opaque gut-feel process into a reviewable ledger.
Planting known traps in simulated resumes before trusting the AI with real ones is a cheap, repeatable validation pattern that applies to any judgment task handed to a model.
The Pending tier is a deliberate design choice that acknowledges the cost of a false negative or false positive is higher than the cost of a 15-minute phone screen.
AI-generated interview questions that cite specific resume lines are harder to game than generic question banks, because they force candidates to defend their own claims.
The unplanned catch of a two-year career gap—something the human designer didn't deliberately plant—shows that rule-based machine checking is more stable than fatigued human attention on repetitive detail tasks.