Test Quality Criteria: Definition
Test quality criteria are the standards used to judge the scientific quality of a diagnostic instrument. The three main criteria work as simple test questions: objectivity – does the result depend on who administers and scores the test? Reliability – does the instrument measure consistently? Validity – does it measure what it claims to measure?
These three questions come from psychological test theory, but they are anything but academic. Anyone choosing an instrument for aptitude diagnostics can use them to examine any offer systematically, from personality questionnaires to cognitive ability tests to interview formats. That is what this article is for: a practical checklist for evaluating an assessment tool without being a psychometrician yourself.
Objectivity: Does the Result Depend on People?
An instrument is objective when the result does not depend on who administers, scores and interprets it. Two candidates with the same abilities should receive the same result, no matter who sits across from them. The literature distinguishes three levels:
- Objectivity of administration: All candidates complete the procedure under identical conditions – the same instructions, the same time limits, the same support. An interview that runs short or long depending on the interviewer’s mood fails this test.
- Objectivity of scoring: The same answers always lead to the same score, regardless of who does the scoring. Standardized instruments with fixed scoring rules meet this almost automatically; free-form impressions do not.
- Objectivity of interpretation: The same score leads to the same conclusion. This requires clear interpretation rules, for example reference values from a norm group.
Digital formats have a structural advantage here: online assessments present every participant with the same tasks and score everyone by the same rules. Objectivity is the easiest of the three criteria to establish – and still where many home-grown selection processes fall short, most often in the unstructured interview.
Reliability: Does the Instrument Measure Consistently?
Reliability describes the precision of a measurement. A tape measure that shows a different value at the same wall every time is useless; the same principle applies to psychological measurement. Every measurement contains some error; the question is how much.
There are three common ways to check reliability, all of them understandable without formulas:
- Test-retest reliability: The same people complete the instrument again after some time. Similar results speak for a dependable instrument – provided the trait itself is stable. Personality should look similar months later; your mood on the day should not.
- Parallel-forms reliability: Two equivalent versions of the same test are completed by the same people. If the results match, the instrument is measuring the trait rather than the quirks of individual items.
- Internal consistency: This checks how uniformly a test’s items capture the same trait. If all questions target the same thing, the answers should form a coherent pattern.
The limit of this criterion matters: a reliable instrument measures precisely, but not necessarily the right thing. A scale that consistently shows too much is perfectly reliable – and still wrong. Precision is a necessary condition, nothing more.
Validity: Does It Measure the Right Thing and Predict Job Success?
Validity asks whether an instrument actually measures the trait it claims to measure, and whether it supports the conclusions you want to draw from it. It is the most demanding of the three criteria. Again, the literature distinguishes three forms:
- Content validity: The tasks directly represent what is being measured. The classic example is the work sample, where applicants perform exactly the activity the job requires.
- Construct validity: The instrument really captures the theoretical trait – conscientiousness, say, rather than the wish to present oneself well. It shows in results that align with related instruments and differ from unrelated ones.
- Criterion validity: Test results are linked to an external criterion – ideally later job success, such as performance ratings, goal achievement or retention. When the criterion is collected in the future, this is called predictive validity.
For personnel selection, criterion validity is the one that counts. It answers the question recruiters actually care about: does this instrument predict who will succeed in the job? Meta-analytic research shows that cognitive ability tests, work samples and structured interviews are among the strongest predictors of job performance, while unstructured interviews and sheer years of experience perform considerably worse (Schmidt & Hunter, 1998).
Validity is never a property of the instrument alone. It is always a statement about a relationship: valid for what, for which population, against which criterion? That is why sound selection does not start with the test but with the requirement profile – only once you know which traits a role really demands can you judge whether an instrument measures the right ones.
The Hierarchy: Each Criterion Builds on the Previous One
The three main quality criteria are not equals; they form a chain.
No reliability without objectivity. If the result depends on who measures, the measuring person becomes a source of error, and the measurement cannot be consistent.
No validity without reliability. An instrument that measures unreliably cannot predict anything reliably. Reliability caps how valid an instrument can ever become.
The chain does not run in reverse: an instrument can be perfectly objective and highly reliable and still measure nothing that matters for the job. Objectivity and reliability are preconditions; validity is the actual purpose. When you evaluate an instrument, never stop at the first two levels – always ask through to validity.
Secondary Criteria: Norms, Fairness, Economy, Acceptance, Utility
Beyond the three main criteria, the literature lists several secondary criteria. They rarely decide a purchase alone, but in practice they often make the difference:
- Norms: Up-to-date reference values from a relevant comparison group exist, so an individual result can be put into context at all.
- Fairness: The instrument does not systematically disadvantage any group, for example by gender, age or background.
- Economy: Effort and cost stand in a reasonable relationship to the insight gained – for the company and for the candidates.
- Acceptance: Applicants experience the procedure as appropriate and understandable. A scientifically strong instrument that drives talent out of the process misses its purpose in practice.
- Utility: The instrument answers a question that would otherwise remain open, and genuinely improves the decision.
Three Questions to Ask Any Provider
Quality criteria become tangible once you translate them into buying questions. Any HR team can ask these three, no statistics required – and they work for every common type of assessment:
- “What validation data can you show us?” A serious provider presents studies or technical documentation explaining how the instrument was tested. A provider who responds with customer logos and testimonials has not answered the question.
- “What sample was it validated on?” How many people, from which occupations, languages and age groups? An instrument validated only on students says little about skilled trade workers – and vice versa.
- “Against which criterion was it validated?” The difference is substantial: correlation with another test? With self-ratings? Or with actual job success, such as performance ratings or retention? For personnel selection, only the last one settles the matter.
A binding framework for this kind of scrutiny is the German standard DIN 33430. It is deliberately a process standard: it does not certify individual tests but defines requirements for the entire process of job-related aptitude assessment – from the requirements analysis through the choice of suitable instruments to the qualification of the people who use them. Quality criteria describe the tool; DIN 33430 describes the professional handling of it. You need both.
Limitations: What Quality Criteria Cannot Do
Quality criteria are the best available framework for judging instruments. Three limitations are still worth knowing.
Validity does not travel automatically. High criterion validity in one study proves the relationship in that sample, with that criterion, in that context. Whether the instrument works equally well in your company, for your target group and your role has not been shown yet. The closer the study context is to your own, the more confidently you can transfer the result – but checking beats assuming.
Quality criteria describe the instrument, not the fit to the role. The best test only measures the trait it was built for. Whether that trait matters for a specific position is decided by the requirements analysis, not by the test manual. A highly valid test of an irrelevant trait improves no hiring decision.
A single figure can look polished and say little. Internal consistency, for instance, can be pushed up by asking very similar questions – which makes the measurement narrower, not better. No single number replaces the full picture: sample, criterion, documentation and how current the norms are.
Test Quality Criteria at Aivy
Aivy is a spin-off of Freie Universität Berlin and builds game-based assessments against exactly these standards: standardized administration and rule-based scoring for objectivity, psychometric testing of every task for reliability, and validation studies against job-related criteria for validity. What that means in practice shows in the Lufthansa success story: a hit rate of 96% in pre-selection. We are happy to answer the three provider questions from this article – sample and criterion included.
Frequently Asked Questions
What are the three main test quality criteria?
Objectivity (the result does not depend on who administers or scores the test), reliability (the instrument measures consistently and precisely) and validity (the instrument measures what it is supposed to measure).
Which quality criterion is the most important?
Validity – in personnel selection specifically criterion validity, because it tells you whether an instrument actually predicts job success. Objectivity and reliability are its preconditions – necessary, but not sufficient.
What is the difference between reliability and validity?
Reliability means an instrument measures precisely. Validity means it measures the right thing. A scale that consistently shows too much is reliable but not valid.
What does criterion validity mean in hiring?
It describes the relationship between a test result and an external criterion, ideally later job success such as performance ratings or retention. It is the form of validity that matters for selection decisions.
Does DIN 33430 replace the quality criteria?
No. DIN 33430 is a process standard: it governs how aptitude assessments are planned, conducted and evaluated professionally. The quality criteria evaluate the individual instrument. A sound selection process needs both.
Sources
- DIN 33430:2016-07: Requirements for proficiency assessment procedures related to occupations. Beuth Verlag.
- Schuler, H. & Kanning, U. P. (Eds.) (2014): Lehrbuch der Personalpsychologie. 3rd edition. Hogrefe.
- Kanning, U. P. (2015): Personalauswahl zwischen Anspruch und Wirklichkeit. Springer.
- Hossiep, R. & Mühlhaus, O. (2015): Personalauswahl und -entwicklung mit Persönlichkeitstests. 2nd edition. Hogrefe.
- Ackerschott, H. et al. (2016): Eignungsdiagnostik – Qualifizierte Personalentscheidungen nach DIN 33430. Beuth Verlag.
- Schmidt, F. L. & Hunter, J. E. (1998): The Validity and Utility of Selection Methods in Personnel Psychology. Psychological Bulletin, 124(2), 262–274.
- Aivy: Lufthansa success story (source for the 96% hit rate in pre-selection).
Make a better pre-selection — even before the first interview
In just a few minutes, Aivy shows you which candidates really fit the role. Beyond resumes based on strengths.




















