Candidate? Explore your strengths

here

Game-based assessments — more than just a gimmick?

Finding your dream job, researching company information and finally submitting an application — these days all of it usually happens on a smartphone. It is no wonder, then, that mobile-optimized selection procedures are becoming increasingly important in candidate selection too (Nikou & Economides, 2018). So-called game-based assessments in particular are growing in popularity.


Gamified or game-based assessments are psychological test procedures presented in a playful format. The idea behind them is that results from psychometrically developed mini-games allow conclusions to be drawn about applicants' cognitive, social and personality-based characteristics. Scientific studies suggest that this is indeed possible (e.g. Brown et al., 2014). Accordingly, more and more large companies such as LinkedIn, Tesla, McKinsey and Deloitte are putting their trust in this new method of candidate selection.

Games on a smartphone instead of multi-hour test batteries and long journeys

That sounds good so far — but are these procedures really suitable for distinguishing between suitable and unsuitable applicants? The question is on many HR managers' minds, and many are still rather skeptical. Justified skepticism, or sleeping through a trend? Time to look at these methods more closely against the key scientific quality criteria for tests. Only when those criteria are met can we speak of a scientifically sound test procedure (Kubinger, 2019).

We will take a closer look at the three main quality criteria of classical test theory (objectivity, reliability, validity) and two important secondary criteria (fairness, economy).

Objectivity

A test procedure is objective when different HR managers arrive at the same assessment of an applicant. Risks to objectivity arise above all in the classic job interview, where the judgment of HR managers can be influenced by all kinds of factors — how likeable they find the applicant, for example. The result is a distorted picture of the applicant's abilities and personality traits, and the procedure is no longer objective. This is where one of the key benefits of game-based assessments comes into play: technology-based delivery and automated scoring, drawing on the latest findings in machine learning, reduce numerous sources of error — objectivity increases. One disadvantage, on the other hand, is that there is no certainty that the person playing is actually the applicant and not someone else (also referred to as impersonation). However, this problem exists in every selection procedure carried out online, and it is one of the reasons why game-based assessments have so far been used primarily in candidate screening.

Reliability

Another important quality criterion is the reliability of a test procedure. Reliability describes the extent to which a test measures a characteristic (e.g. a cognitive ability or a personality trait) accurately, that is, without measurement error. As a rule, this works better the more data is available about the person. A simple example makes this clear: imagine taking a concentration test after a long, stressful day at work, disturbed by your neighbors' loud music. The result will probably not reflect your actual ability to concentrate (a so-called trait), but rather a situational, temporary measurement of it (a so-called state, cf. Fleeson, 2001). You will certainly perform worse than on a day when you are well rested and undisturbed. If your ability to concentrate is measured across several days, however, such random measurement errors increasingly cancel each other out — in statistics this is known as the central limit theorem. This is exactly where the idea behind game-based assessments comes in. Instead of inferring an applicant's ability to concentrate from a single test result, the results of several game runs are stored and averaged. The outcome is a more accurate estimate of the applicant's actual ability to concentrate.

Validity

But of course a procedure should not only measure precisely — it should also measure the right thing (validity). In candidate selection, the primary interest is predictive validity: the test should predict a particular outcome as accurately as possible. A commonly used outcome is the applicant's future job performance. As with reliability and objectivity, various sources of error can affect predictive validity. One is applicants' response tendencies. A particular problem with classic test methods is social desirability — the tendency of applicants to deliberately choose answers that present them in a positive light. This matters especially when the purpose of the test is easy to see through, which is frequently the case in personality assessment. Take a statement from the Big Five, one of the best-known personality instruments (Asendorpf & Neyer, 2012). Applicants are asked how far they agree with the following:

“I see myself as someone who is reliable and conscientious.”

Clearly, very few people would disagree with such a statement while applying for their dream job. Whether the item is suitable for distinguishing between conscientious and less conscientious applicants is therefore questionable.


Applicants do not always actively try to manipulate results, though. Just as often, they simply find it hard to judge their own personality traits, strengths and weaknesses. Because what does it actually mean to be conscientious or extroverted? And how conscientious or extroverted am I really? That is often not easy to answer, and psychology refers to it as a limited capacity for introspection. To simplify the question, applicants tend to compare themselves with the people around them. What they end up answering is: how conscientious or extroverted am I compared with the people in my environment? A range of scientific studies shows that this shift in the question frequently produces distortions (e.g. Schwarz, 1999).


Game-based assessments sidestep exactly this problem

Instead of relying solely on the applicant's self-report, the mini-games additionally capture nuances of behavior. The applicant's preference for speed versus accuracy, for example, is observed across the individual games. Those behavioral nuances are then used to complement the error-prone self-report with objective data. Weighting the data optimally is achieved with the help of intelligent self-learning algorithms. Capturing actual behavior instead of relying on what the applicant says about themselves — sounds logical, doesn't it? And scientific research confirms that for many characteristics this approach produces more valid results than self-report alone (e.g. Baumeister et al., 2007).

Test fairness

Another important quality criterion is test fairness: no group should be systematically disadvantaged by a test procedure, for example on the basis of gender or ethnic background. With many test methods this is nonetheless the case, because the questions are geared towards Western cultures (Camilli, 2006). Gender, ethnic background and skin color, by contrast, play no role in the results of the largely language-free mini-games. Nor is there any scientific evidence to date for the occasionally voiced concern that people with gaming experience might have an advantage. Assessment procedures evaluate only the factors relevant to job success and leave irrelevant characteristics such as gender or social background out of the picture. This creates greater fairness and equal opportunity.

Test economy

The final criterion, test economy, was touched on at the outset. Selection procedures should be characterized by a short duration, low costs and little effort for the applicant. Game-based assessments follow the idea of a “zero-footprint” measurement — the burden on applicants all but disappears, and the games are often genuinely fun. Because the individual tasks can be adapted to the applicant with the help of intelligent algorithms (known as adaptive testing), the games do not become boring either. And it is not only applicants who benefit. The cost and time savings compared with other test methods (e.g. an assessment center) are enormous — particularly with very heterogeneous and international target groups.

Download now for free: our tabular overview of the key sources of error in occupational aptitude assessment and how game-based assessments address them.

Conclusion

The analysis shows that these modern procedures can hold their own against established psychometric methods. Fun, measurement accuracy, validity and fairness are not mutually exclusive. In combination, they contribute to a candidate selection process that follows scientific standards and meets the changing needs of a new target group.

Implementing all quality criteria in practice is not always easy, however. Many psychometric games on the market do not meet the criteria described above (König et al., 2010). The procedures should therefore be reviewed critically before being used as part of a professional selection process. DIN 33430 on occupational aptitude assessment provides an important frame of reference here.

Your assistant for talent assessment

Try it for free

Become a HeRo 🦸 and understand candidate fit - even before the first job interview...