The Doors We Build With Scores
Psychometrics, present-day AI and the intelligence still to come
A score is a small thing. A number on a page. A percentile on a screen. A mark above or below a line. Yet institutions build doors around it.
One score leads towards assistance; another towards exclusion. A child is offered support or placed in a different educational setting. An applicant reaches an interview or disappears from the process. An employee becomes eligible for promotion. A student enters a university. A patient receives a diagnosis that may change how others understand them—and how they understand themselves. The score passes from the language of measurement into the language of action.
This is why psychometrics should never be regarded as merely a branch of mathematics. It is an applied science whose findings enter schools, workplaces, hospitals, courts and systems of government. Its methods may be statistical, but its consequences are human. For more than a century, arguments about reliability, validity, standardisation, fairness and diversity have therefore concerned more than the technical quality of tests. They have concerned the authority that institutions may legitimately exercise over people.
The central question is not only: What does this score tell us?
It is also: What are we entitled to do because of it?
That question reveals a recurring mismatch. Psychometrics has become steadily more exact in the way it measures people, while institutions have often remained blunt in the way they act upon the result. A finely estimated score reaches a threshold, and an entire future is divided: pass or fail, admit or reject, support or exclusion.
The measurement becomes more discriminating.
The decision remains a rubber stamp.
The authority of the number
A psychometric test score carries a particular kind of authority. It may be expressed precisely, compared with a standardisation sample and accompanied by an estimate of reliability or uncertainty. It therefore appears quite different from an ordinary human opinion.
But the number cannot interpret itself. Someone has decided what should be measured. Someone has chosen which behaviours will count as evidence. Someone has selected the comparison group. Someone must decide what a high or low score means. Most importantly, someone must decide what should follow.
Reliability can tell us whether a score is sufficiently consistent to support interpretation. It cannot tell us whether the interpretation is relevant or just. Standardisation can tell us how a person compares with others under specified conditions. It cannot tell us which opportunities that comparison should open or close. Validity concerns the evidence supporting particular interpretations and uses of test scores. It does not transform every technically defensible inference into a legitimate institutional action.
Professional testing standards accordingly treat fairness, accessibility, score interpretation and the consequences of test use as central rather than peripheral concerns. Psychometric evidence may contribute to a judgement. It should never be mistaken for the whole of the judgement.
Binet and the possibility of help
In 1905, Alfred Binet and Théodore Simon published the first version of their intelligence scale, intended to help identify children who might need additional educational support. They resisted the belief that the resulting classification revealed a fixed and unalterable quantity of intelligence.
But classifications do not always remain attached to the purposes for which they were created. A description intended to direct assistance can become a label. The label can become an explanation of the person. The explanation can then become a prediction of what that person will never be able to do. The movement is gradual.
The test begins by saying: This child appears to need help with these tasks.
The institution begins to hear: This is the kind of child this is.
A provisional judgement becomes an identity. The danger lies not only in inaccurate measurement. It lies in converting limited evidence into excessive authority. The score may establish something modest. The institution may claim something far greater. The score has not changed. Its purpose has.
When the gateway becomes the judgement
In 2019, the same distinction appeared in a contemporary employment case. Kevin Meier was a computer-science graduate who applied for a position with British Telecommunications. He had autism, dyslexia and dyspraxia. Before he could reach an interview, he was required to pass an online situational judgement test. His performance prevented him from progressing.
The Court of Appeal in Northern Ireland upheld the finding that BT had failed to make a reasonable adjustment. The company could have allowed him to bypass the test or declined to treat its result as an absolute barrier to an interview. The case did not establish that situational assessments are inherently unjust. It revealed the danger of allowing one procedure, in one format, to become the entire judgement.
The test had been designed to provide evidence about suitability. Instead, it was used to prevent any other evidence from being considered. A procedure intended to assess judgement had itself been applied without sufficient judgement.
When a pattern appears before its explanation
In Essop v Home Office, Mr Essop, an immigration officer, led a group of 49 current or former Home Office employees who challenged the requirement to pass a Core Skills Assessment before becoming eligible for promotion. They argued that the requirement amounted to indirect discrimination on grounds of race and age because candidates from ethnic minorities and older candidates passed at substantially lower rates than white and younger candidates. The precise causes of those differences had not been established.
In its judgment of 5 April 2017, the Supreme Court held that the claimants did not first have to explain why the assessment produced the group disadvantage before it could become legally relevant. Once an apparently neutral requirement was shown to place a protected group—and the individual claimant—at a particular disadvantage, it was for the employer to justify that requirement as a proportionate means of achieving a legitimate aim.
Ethical responsibility does not begin only when the explanation is complete. Nor can it be suspended because asking further questions is inconvenient. Uncertainty is not a licence to look away. It is a reason to look more closely.
An examination that never happened
When public examinations were cancelled in the UK in 2020 because of the Covid-19 pandemic, teachers estimated the grades their pupils might have achieved, while examination authorities attempted to preserve standards between years through statistical adjustment.
Everyone was accustomed to waiting for the examination boards to decide. But there were no examinations for them to assess. Standardisation could normally compensate for differences between papers pupils had actually sat. It could not recover performances that had never occurred. An educational institution’s history might suggest how many high grades were probable, but not which pupils would have earned them. The error was not simply choosing the wrong algorithm. It was asking a model to provide an answer that its evidence could not support.
Alternative arrangements—accepting reduced comparability, relying more heavily on teacher judgement, holding later examinations or changing university admissions—needed to be considered as policy choices. No solution would have been perfect. But the uncertainty should have been acknowledged and shared, rather than concealed inside a procedure that made an impossible judgement appear statistically authoritative. Evidence subsequently given to the UK Covid-19 Inquiry by Sir Jon Coles exposed the same category error: a model capable of preserving the national distribution of grades was treated as though it could determine which individual pupils would have earned them.
The procedure could reproduce a pattern.
It could not recreate the missing event.
When a game becomes a gatekeeper
A growing number of employers now use game-based assessments in recruitment. Applicants complete apparently simple tasks on an app, while patterns of speed, persistence, risk-taking and response to reward are translated into estimates of personality and suitability for work. The applicant may not know what is being measured or why the resulting profile has closed the door to further consideration. The employer may understand little more.
In 2026, Rishi Bommasani and colleagues analysed more than four million applications from 3.4 million people across 156 employers, all assessed through algorithms supplied by a single vendor. Yet even with this rare access, the researchers could not answer the central psychometric question: did the assessments actually identify applicants who were better suited to the jobs? They had no independent measures of applicant quality or job performance with which to test the validity of the models.
The disturbing point is the scale at which such procedures can become accepted before their effectiveness has been independently demonstrated. They appear impartial because they are standardised, and authoritative because they are marketed as scientific and objective. An uncertain judgement can become a gatekeeper for millions of applicants.
When the question disappears
Across these examples, the technology changes but the same movement recurs. Evidence is produced for a limited purpose. The evidence enters a procedure. The procedure becomes routine. The routine acquires authority. Then the original question disappears.
Binet’s assessment no longer asks what assistance a child may need. It becomes a statement about what kind of child this is. An employment test no longer contributes one piece of evidence about suitability. It prevents any further evidence from entering. A statistical model no longer acknowledges what cannot be known. It produces the appearance of an answer. An unvalidated game ceases to be an experimental instrument and becomes a gatekeeper.
The danger lies not only in closing the door. It lies in closing the enquiry. Once a threshold has been embedded in an institutional process, it begins to appear natural. The distinction it creates looks as though it had always existed. A decision made by people becomes mistaken for a property of the score. Those operating the system may eventually forget that someone chose where the line should be drawn.
Bringing judgement back in
The obvious response is that human beings must be brought back into the decision. That response is right. A consequential judgement should not be surrendered to a score, an algorithm or an automated system. Someone must remain responsible for deciding what the evidence means, what may have been missed and whether the consequences are justified. Where a door is opened or closed, human judgement should not disappear behind the machinery.
But bringing human beings back does not mean returning to unaided intuition. Human judgement is not naturally fair, consistent or transparent. It is influenced by confidence, familiarity, social resemblance, personal preference and assumptions that may never be made explicit. An interviewer may be impressed by qualities unrelated to later performance. A teacher may underestimate a pupil who does not resemble previous successful pupils. A manager may interpret difference as deficiency without recognising that a judgement has been made.
Psychometrics developed partly to discipline such decisions. A well-constructed assessment can organise evidence, make comparisons more consistent and expose institutional claims to examination. It can quantify uncertainty, reveal differences between groups and test whether confidence is supported by results. Properly used, it does not replace human judgement. It gives human judgement evidence against which to test itself.
Institutions also face unavoidable problems of scale. If a thousand people apply for twenty positions, not everyone can be interviewed. If educational or clinical resources are limited, some method of allocation must be found. Scores can help make judgement possible at scale. But they should narrow an enquiry, not end it.
The failure begins when psychometric evidence—through institutional convenience, commercial pressure or professional overconfidence—is allowed to become the decision itself. A score may inform a judgement. It cannot assume responsibility for one.
Could AI help?
If human judgement must be restored, perhaps AI could help humans exercise it more carefully. An AI system could bring together evidence that no individual decision-maker has time to examine. It could identify inconsistencies, point out missing information, compare competing interpretations and make uncertainty visible. It might challenge an interviewer’s first impression, detect when a score conflicts with other evidence or ask whether the criteria being applied are appropriate to the decision being made. Used in this way, AI would not replace human judgement. It would enlarge the evidence available to it.
Yet this suggestion may inspire little confidence. By this point, the reader may feel that institutions have already surrendered too much authority to systems they do not fully understand. Adding AI may seem not like the restoration of judgement, but its final removal: one more layer of machinery between the person affected and the human being who ought to take responsibility.
That concern is justified. AI is not an answer if it merely produces another score, another ranking or another recommendation that no one feels able to question. It becomes useful only when it helps human beings deliberate, disagree and remain answerable for what follows. The human beings must remain. The question is whether AI could help them deliberate without quietly taking the decision away from them. A single AI adviser would offer little protection. Its recommendation might acquire the same misplaced authority as a score, while making the reasoning behind it still harder to challenge.
Nor is the difficulty solved merely by multiplying AI agents and allowing them to deliberate in place of human beings. One agent may represent the evidence, another fairness, another legality and another doubt. Their exchange may appear richer than the output of a single system, and it may genuinely help human decision-makers see what they would otherwise miss. But many agents can still inherit the same assumptions, serve the same institutional purpose and converge upon a destination chosen before their deliberation began. A chorus of artificial voices does not amount to independent intelligence if none can question why the song was chosen.
The limitation is therefore deeper than scale, fluency or number. Present systems may be able to describe an ethical objection without possessing any power to make it consequential. They may recognise that a threshold is poorly justified and still be required to apply it. They may identify a pattern of disadvantage while remaining embedded in the process that produces it.
This is why human responsibility cannot simply be transferred to a more elaborate assembly of machines. AI agents may widen the evidence, expose disagreement and challenge assumptions. But someone must still be able to reconsider the purposes being served, alter the criteria being applied and refuse the destination towards which the system is moving. The question is not simply whether AI can speak the language of ethics. It is whether the deliberative process—human, and perhaps one day partly artificial—can reconsider the direction of the system itself.
The intelligence that is missing
This may require something closer to genuine artificial general intelligence. AGI is often imagined simply as a system capable of performing a wider range of human tasks. But breadth of competence would not be enough. A machine able to do more things might still pursue each assigned purpose without examining why it had been assigned.
The intelligence needed here would have to be reflective as well as general. It would need to recognise when the evidence no longer supports the authority being claimed for it. It would need to notice when a practical rule had quietly become a moral principle, when administrative convenience had replaced the purpose of the assessment, or when the question being answered was no longer the question that ought to have been asked. It would have to bring the task itself into the enquiry:
What are we doing?
Why are we doing it?
What has this score been allowed to become?
Should the direction now change?
This does not mean refusing to decide. An enquiry that never closes is no more useful than one that closes too soon. Institutions must act, and some people will still be excluded. The challenge is to reach a decision without allowing the decision to erase the question from which it arose.
My work on human–AI and AI–AI trajectories begins with that problem: whether an enquiry can remain capable of changing direction without becoming incapable of conclusion. Existing AI agents may help us explore it. They may reveal assumptions, introduce competing perspectives and expose moments of premature closure. But they should not be mistaken for its solution. More precise scores will not solve the post-score problem. More powerful models will not necessarily solve it. More agents will not solve it if they remain confined within the same inherited purpose.
Psychometrics can discipline the evidence on which institutional decisions are based. Artificial intelligence can extend our ability to examine that evidence. But to take the matter further, we may need an intelligence capable not merely of answering the question placed before it, but of recognising when the question itself must be reopened.
Institutions will continue to build doors around numbers. They have to act. They have to choose. Some doors will open and others will remain closed. A score becomes dangerous not merely when it closes a door. It becomes dangerous when it closes the question of whether the door belongs there.
A score is a small thing.
True intelligence begins when the door can no longer make us forget that it was built.
© John Rust, July 2026, All Rights Reserved
References
AERA, APA and NCME, Standards for Educational and Psychological Testing (2014). The Standards explicitly connect validity to intended score interpretations and uses, and treat fairness as foundational.
Binet and Simon’s original work on diagnosing the intellectual level of children in schools and institutions.
Essop v Home Office [2017] UKSC 27.
Ofqual, An Evaluation of Centre Assessment Grades from Summer 2020 (2021).
Bommasani et al., “Algorithmic Monocultures in Hiring,” FAccT 2026.
Coles, J. (2025). Oral evidence to the UK Covid-19 Inquiry, Module 8: Children and Young People, Day 5, 6 October 2025, pp. 68–74 and 89–90. UK Covid-19 Inquiry.
Rust, J. (2025) Teleosynthesis: The Direction in the Algorithm


