Work Sample Test: Definition
A work sample test is a standardized task used in personnel selection. It mirrors a representative slice of the actual job and is scored against criteria defined in advance. Candidates demonstrate real work behavior instead of merely talking about it.
The principle: to test complaint handling, have candidates answer an anonymized complaint; to test coding, set a small coding task. The task simulates the job – under identical conditions for everyone, with a scoring rubric fixed before the first candidate sees it. These two features, standardization and predefined criteria, separate a work sample test from an improvised “show us what you can do”.
Within aptitude diagnostics, work sample tests belong to the simulation-oriented class of methods: they measure what a person can do. Biography-oriented methods ask what someone has done in the past, while construct-oriented methods such as the cognitive ability test capture the stable traits a person brings along (Schuler & Kanning 2014).
Work sample, trial workday, internship: three different things
In day-to-day recruiting, the three terms are often used interchangeably. Methodologically and legally, the distinction matters.
The work sample test is a practice task. It simulates the job, takes minutes rather than hours, and its output is never put to productive use. Every candidate works on the same task under the same conditions – otherwise the results are not comparable.
The trial workday (in Germany: Probearbeit or Schnuppertag) takes place in the real business. Under German employment law it is meant as a non-binding get-to-know arrangement: observing, shadowing, tagging along – without an obligation to work. As soon as candidates deliver real, usable work, a claim to remuneration typically arises (Section 612 of the German Civil Code, BGB). And methodologically, a trial day can hardly be standardized: every candidate experiences a different day – without a clear rubric, all you get is a second unstructured impression.
The internship is a separate legal relationship with a learning purpose, stretching over weeks or months. It is neither designed nor suited as a selection instrument for a specific role.
A second common mix-up concerns work samples that candidates bring along: portfolios, writing samples, reference projects. These document past work and belong to the biography-oriented class. They are valuable – but they answer a different question: what did this person produce in the past, under unknown conditions and with unknown support?
Why work sample tests are among the most accurate methods
According to the current, conservative estimates by Sackett et al. (2022, 2023), work sample tests rank among the strongest single predictors in personnel selection at r ≈ .33 – behind the structured interview (r ≈ .42) and roughly on par with cognitive ability tests (r ≈ .31). Older analyses reported considerably higher values; they rested on very few primary studies and were corrected downward by broader meta-analyses (Hunter & Hunter 1984; Roth, Bobko & McFarland 2005).
Three properties explain the method’s strength.
High content validity. A work sample test directly measures the behavior the job will demand – not a construct from which that behavior first has to be inferred. The more representative the task, the shorter the inferential chain between test result and job performance.
High face validity and acceptance. Candidates immediately see why the task is set. That is more than cosmetics: methods with a visible link to the job are perceived as fairer – Hausknecht et al. (2004) find work samples and interviews among the best-accepted selection methods.
Little room for self-presentation. A work sample rewards what someone does, not how skillfully someone talks about themselves. Background, educational path and rhetorical talent lose weight; demonstrated behavior gains it.
The assessment center, often treated as the gold standard, bundles several such simulations – yet a single, carefully built work sample test achieves predictive power in a similar range, at a fraction of the effort (Sackett et al. 2023).
How to build a good work sample test: five steps
A work sample test is quick to invent and hard to do well. The difference lies in the preparation.
- Requirement analysis first. Start from the requirement profile: which activities form the core of the role, and which decide between success and failure? Without this, the work sample measures what is easy to test – not what matters.
- Choose a representative task. Ideally a real, anonymized task from everyday work: a customer complaint, a buggy piece of code, an inbox that needs prioritizing. The task should capture the critical slice of the job and be solvable in a short time.
- Standardize the administration. Same task, same time limit, same materials, same instructions – written down and identical for every candidate. DIN 33430 provides the quality framework for job-related aptitude assessment.
- Anchor the rating scale. Three to four criteria are enough, but each needs behavioral anchors: a short description of which concrete behavior corresponds to which scale level – written before the first administration, not after it.
- Use several independent raters. At least two people score separately and only compare their judgments afterwards. This dampens leniency and severity tendencies and reveals where criteria are read differently.
An example from customer service: a real, anonymized complaint, to be answered in writing within a set time, scored on tone, solution quality and prioritization. Built once in a few hours – after that, every administration produces comparable, documented results.
Common mistakes in practice
Unpaid work on the real product. Having candidates solve a live, open task and then using the result puts you on thin ice legally (Section 612 BGB) and damages trust. The task should be anonymized, already solved internally, or clearly fictitious.
Excessive time demands. Take-home assignments running for hours or days shift the cost of the process onto candidates. In-demand candidates drop out first – and the test ends up selecting for spare time rather than competence.
Unclear criteria. A standardized task with freehand scoring is pseudo-structure: it creates the appearance of objectivity while the judgments remain gut feeling. Without an anchored scale, there is no work sample test in the diagnostic sense.
Scoring by committee. If raters discuss their impressions before each person has judged independently, the loudest voice wins – not the best observation.
Tasks without a link to the job. Brainteasers and trivia quizzes feel demanding but represent no slice of the actual work – removing exactly the mechanism that makes work samples accurate, and the perceived fairness with it.
Limits of the work sample test
Anyone using the method should know four limits.
It measures competence, not potential. A work sample shows what someone can do today. Candidates who have never performed the activity will fail the task – no matter how quickly they would learn it. For career changers, entry-level hires and apprenticeships it is therefore unsuitable as the main criterion; cognitive ability and learning potential tell you more there.
It is effortful. Construction, administration and scoring take people and time. Early in the process, with high applicant volumes, that barely scales – the work sample belongs in the later stages, once the field has been narrowed down.
It is a snapshot. Daily form, nerves and the test situation feed into the result. And it measures a narrow slice: how someone collaborates over months, handles feedback or learns new topics is not visible in a single task.
It ages with the job. When tasks, tools or processes change, the work sample has to follow – otherwise it measures yesterday’s role.
Combining work samples wisely
No single method should carry a hiring decision on its own. The strength lies in the combination: a structured interview, a work sample test and a construct-oriented test each capture information the others cannot see – accuracy rises with every additional class of methods (Schuler & Kanning 2014; Sackett et al. 2023).
This is exactly where Aivy picks up what work samples leave open: the game-based, scientifically validated assessments of the Freie Universität Berlin spin-off measure cognitive abilities and personality traits – potential, in other words – early in the process and at scale, long before a work sample becomes practical. More than 1 million completed assessments across 200+ companies show how well the two approaches complement each other: the assessment shows who brings the potential for the role; the work sample shows, later in the process, who already masters the competence.
Frequently asked questions
Is a work sample test the same as a trial workday?
No. The work sample test is a short, standardized practice task with fixed scoring criteria. The trial workday takes place in the real business, can hardly be standardized and is legally intended as a non-binding get-to-know arrangement.
Do candidates have to be paid for a work sample test?
Not for a short, simulated practice task without usable output. Once real, usable work is delivered, a remuneration claim typically arises under German law (Section 612 BGB). When in doubt: anonymize the task or use one that has already been solved internally.
How long should a work sample test take?
As short as possible: a good work sample is measured in minutes, not days. What matters is that the critical slice of the job becomes visible – extra length rarely adds insight but increases drop-out.
Do you need a certification to use work sample tests?
No. Unlike standardized psychometric tests, work samples are license-free and can be developed and administered in-house with careful preparation. DIN 33430 offers the quality framework – from requirement analysis to documentation.
Which roles are work sample tests suited for?
Anywhere the core of the job can be captured as observable behavior in a short task: customer service, software development, editorial work, sales, skilled trades. Less suitable where candidates cannot yet master the activity – as in apprenticeships and career-changer hiring.
Sources
- Sackett, P. R., Zhang, C., Berry, C. M. & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Sackett, P. R., Zhang, C., Berry, C. M. & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283–300.
- Roth, P. L., Bobko, P. & McFarland, L. A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology, 58(4), 1009–1037.
- Hunter, J. E. & Hunter, R. F. (1984). Validity and utility of alternative predictors of job performance. Psychological Bulletin, 96(1), 72–98.
- Hausknecht, J. P., Day, D. V. & Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57(3), 639–683.
- Schuler, H. & Kanning, U. P. (Eds.) (2014). Lehrbuch der Personalpsychologie (3rd ed.). Hogrefe.
- Kanning, U. P. (2015). Personalauswahl zwischen Anspruch und Wirklichkeit. Springer.
- DIN 33430:2016. Requirements for proficiency assessment procedures and their implementation. Beuth.
Make a better pre-selection — even before the first interview
In just a few minutes, Aivy shows you which candidates really fit the role. Beyond resumes based on strengths.




















