Transparency
How we score
Every score on SkillChecker is produced by a defined, consistent process, not an opaque one. This page explains exactly how it works, and what it does not claim.
Why we use AI
Human reviewers grading dozens of answers tend to drift — the same answer can score differently depending on what was read just before it, or an unconscious preference for how something is phrased. SkillChecker uses a large language model from OpenAI to grade every answer against the same written criteria, with no name, resume, or photo in front of it — just the question and the answer. (We don't pin this page to a specific model version, since the underlying model may change over time — the criteria below are what stays constant.)
That prevents one specific failure mode — favoring a participant because of who they are — though it doesn't mean the model's own judgment is free of bias in every other respect. The next section explains what that does and doesn't cover.
Bias and its limits
A fair question is whether the model favors certain writing styles, dialects, vocabulary, or levels of formality, even without knowing who wrote an answer. We can't honestly claim it doesn't — no one building on top of a large language model can responsibly claim that. These models are trained on huge amounts of text and can carry subtle preferences toward particular phrasings or registers that have nothing to do with whether an answer is correct.
What we can point to is a real, specific mitigation, not a promise: every grading prompt explicitly instructs the model to score based on accuracy, depth, and practical understanding — and explicitly not on writing style, length, or polish. That constraint is genuinely present in the prompt, not just a claim about it, and it's designed to reduce the risk of confident or eloquent phrasing being mistaken for correctness. We haven't run a formal study quantifying that reduction, so we won't claim one — it does not guarantee the model has no residual bias, which is a limitation of the underlying technology, not something this product has solved.
This is also why we treat a score as an input to a decision, not the decision itself. Every report comes with follow-up interview questions specifically so a human can probe further before anyone is ruled in or out.
What your score is compared against
SkillChecker runs two different kinds of assessment, and they check different things:
- Skill verification — checks a participant's answers against their own claimed skills and experience, nothing else. The question "does this answer demonstrate the skill this specific person claims to have" is what's being scored.
- Role assessment — checks every participant against a fixed baseline for the role itself, with no individual's claims involved at all. Everyone screened for a given role is measured against the same standard, which is what makes their scores directly comparable to each other.
How each answer is graded
Every answer is graded primarily on accuracy, depth, and practical understanding of the specific skill the question targets — reasoning quality and judgment are part of that: under our rubric, a shallow or illogical answer scores lower on depth and understanding, not automatically, but because that's what those two criteria are specifically designed to capture. What's explicitly excluded is writing polish, length, and formatting: a short, precise, correct answer scores higher than a long, vague one dressed up in confident language. The model is specifically instructed to be skeptical of generic, hand-wavy answers that avoid specifics, since that's treated as a sign the underlying claim may not hold up.
One honest exception: if a role genuinely requires strong written communication, that has to be built into the question itself (e.g. "write the message you'd send to X"), not assumed automatically — SkillChecker doesn't separately score communication quality unless a question is designed to test it directly.
Every answer is also graded in context of how it was given. An assessment taken live, in the middle of an interview with no time to prepare, is graded for correct reasoning and direction — not polish or completeness. One completed beforehand, at the participant's own pace, is evaluated against the full expectations of the assessment. And if a participant's timer ran out before they finished, short or missing answers caused by that aren't held against them — only the substance of what they wrote is judged.
Where the overall score comes from
Each question gets its own score from 0–100, with a sentence or two of specific feedback explaining why. The overall score shown on every report, ring, and dashboard is a plain average of those per-question scores — every question weighted equally, computed by our own code, not a separate judgment the AI makes up on its own. If you can see the per-question scores, you can recalculate the overall score yourself.
Should the same answer score exactly the same twice? We deliberately turn down the model's randomness for grading, so a repeat should land in the same range, not swing wildly — but we won't claim it's identical down to the last point every time. Language models are not deterministic calculators; a small amount of run-to-run variation is an inherent characteristic of the technology, not a flaw specific to this product. What stays fixed is the criteria being applied, not necessarily the exact number.
How a verdict is decided
The label you see — Verified, Partially Verified, or Unverified (Strong/Adequate/Weak Baseline for role assessments) — is not the AI's own opinion. It comes from one fixed rule, applied the same way to every report, everywhere in the product:
| Overall score | Verdict |
|---|---|
| 75–100 | Verified / strong baseline |
| 50–74 | Partially verified / adequate baseline |
| 0–49 | Unverified / weak baseline |
Why these particular cutoffs, and not some other split? They're a deliberate design choice, not a number derived from a statistical validation study — we haven't run one, and we'd rather say so than imply otherwise. The reasoning behind them: 75 sets the threshold for "Verified" above the midpoint on purpose, so an average performance doesn't count as a pass — it takes genuinely strong answers across the board, not just barely-adequate ones. Below 50 means the average across their answers fell short of the required threshold. The same two numbers are used everywhere in the product, so the ring's color, the badge text, and the dashboard verdict can never disagree with each other.
What a score isn't
Not comparable across different assessments. A 75 on one role's screening and a 75 on a different role's screening aren't measuring the same thing — different questions, different skills, different required depth. Scores are only directly comparable to other scores from the same screening, or to the claims made on the same CV Verification.
Not a measure of the whole participant. It reflects how a specific set of answers, given at one point in time, held up against this rubric — not intelligence, potential, or overall fit for a role. Plenty of real signals about a participant sit outside what a written assessment can capture.
Not a percentage of "correct" answers. Unlike a multiple-choice test, there's no fixed answer key being checked off — each answer gets a holistic 0–100 judgment against the criteria above, and the overall score is the average of those judgments, not a tally of right versus wrong.
What this doesn't do
This tool is built to support a hiring decision, not replace one. It doesn't verify identity, and it doesn't detect whether someone had outside help while answering. Treat a report as one well-structured piece of evidence, not a verdict on a person.
Answer integrity signals
SkillChecker flags three things while a participant is taking an assessment: whether a paste event occurred, whether text appeared on screen faster than a person can actually type, and whether they left the browser tab multiple times. All three are timing and browser-event facts, not a guess about writing style — we deliberately don't try to judge whether an answer "looks AI-written," because that approach is unreliable and is well-documented to misfire against non-native English speakers and other legitimate writing styles far more often than it catches anything real.
A single tab switch isn't flagged on its own — that's common and often innocuous (checking the time, a notification). It only becomes a signal once it happens more than once.
None of these flags change a score. They're shown to the recruiter as a prompt to ask a follow-up question, not an automated accusation — pasting isn't proof of anything on its own; someone might paste in notes they wrote themselves earlier, and leaving the tab could just as easily mean a notification came in.
Where this sits under the EU AI Act
Using AI to evaluate candidates for a job is treated as a high-risk use under the EU AI Act. We are not going to argue our way out of that category — assessing people for work is exactly what it was written for, and this tool assesses people for work.
Some of what the Act asks for is already on this page, because it was the right thing to publish regardless: how scoring works, the thresholds, what we deliberately exclude, and where the limits are. Some of it is behind the page — a person at the hiring company decides, they see the answers and not just a number, integrity signals never touch the score, and answers are deleted on a published schedule. The Act also prohibits inferring emotion from people in a recruitment setting. We do not do that, and there is no camera or microphone involved at any point.
What we are not doing is claiming to be finished. Our formal classification under the Act, and the paperwork that follows from it, is with our legal advisers. We would rather say that plainly than display a badge we have not earned.