GradeLab is building the AI operating system for assessments. We help schools, universities, and organisations automate handwritten exam evaluation using OCR, large language models, and intelligent grading pipelines, so every student gets a fair, consistent, and accurate result.
We're looking for a QA Engineer who thinks like both an attacker and an end user. This isn't a traditional QA role: you'll validate AI behaviour, grading accuracy, OCR quality, and system reliability across a fast-moving production platform.
What you'll do
• AI evaluation and benchmarking — build golden datasets from real exam papers, measure grading accuracy against human evaluators, and catch regressions before they ship.
• AI observability — use Langfuse and evaluation frameworks to monitor grading accuracy, human agreement, OCR confidence, rubric compliance, cost, and latency in production.
• OCR validation — stress-test the pipeline against poor handwriting, multiple languages, low-resolution scans, folded papers, diagrams, equations, and every real-world document quirk you can find.
• Grading validation — test partial marking, step-wise marking, rubric compliance, and score calculations for consistency between human and AI evaluators.
• Prompt evaluation and AI safety — run adversarial testing including prompt injection, jailbreak attempts, and context poisoning to keep the system secure.
• Edge case testing — break the system with real-world scenarios: crossed-out answers, missing pages, torn sheets, blank pages, and everything students actually submit.