Insights Whitepaper
Whitepaper

When Written Fluency Stopped Being a Signal

Adeolu Timothy 25 February 2026 9 min read Selection Integrity · Governed AI · Assessment Design · Programme Design · Applied AI
Summary

Fellowships, scholarships and grant programmes screen at scale on written English. Generative tools made that filter measure access rather than readiness, and detection cannot fix it: a Stanford study found detectors misclassify more than half of non-native TOEFL essays as machine-written while scoring near-perfectly on US eighth-grade work. This is the case for replacing the essay as a screening instrument, and an account of what we built in its place for a programme where admission carries real money.

Summary. Written English is the cheapest way to screen thousands of applicants, which is why almost every fellowship, scholarship and grant programme leans on it. That filter has quietly stopped working. It now measures who had access to a competent writing tool. Detection is not a remedy, because detectors fail hardest on exactly the applicants these programmes exist to reach. This document sets out what the essay was actually a proxy for, what can replace it, and what it cost us to build an alternative for a programme where admission carries stipends and mentorship rather than just a certificate.

#The filter that quietly stopped working

Selection at scale is an economics problem before it is a fairness problem. A UN agency call for proposals, a Chevening or Fulbright round, an Erasmus Mundus or DAAD intake, a Commonwealth Scholarship cycle, a Gates Foundation grant window, a Y Combinator batch: all of them face far more applicants than they can interview. Something has to reduce the pile before humans look closely, and for decades the written statement did that work cheaply and defensibly.

The statement was never valued for its prose. It was a proxy. Panels read it as evidence of clarity of thought, seriousness of intent, ability to structure an argument, and enough command of the working language to function in the programme. Prose quality correlated with those things well enough to be useful.

That correlation has broken. By 2024, roughly half of surveyed college applicants were using AI to brainstorm essays, 47 per cent to build an outline, and about a fifth to generate first drafts (Boston Globe, Hechinger Report). Among US teenagers, the share using ChatGPT for schoolwork doubled between 2023 and 2024 (Pew Research Center).

Institutions have not caught up. A 2025 Kaplan survey of admissions officers found that 68 per cent of colleges had no policy at all on generative AI in admissions essays. Two per cent explicitly permitted it, 30 per cent banned it (Kaplan).

When a signal is available to everyone at no cost, it stops being a signal. It becomes a formality that consumes reviewer attention without discriminating between candidates.

The practical effect is worse than noise. A polished statement used to be weak evidence of capability. It is now reliable evidence of one thing only: that the applicant had a device, a connection, and the confidence to use a tool. Those correlate with privilege. The filter has inverted.

#Detection is not the remedy

The instinct is to detect and disqualify. This is the wrong move, and the evidence against it is unusually clear.

Liang and colleagues at Stanford tested widely used GPT detectors against writing from native and non-native English speakers. The detectors misclassified more than half of non-native-authored TOEFL essays as AI-generated, while achieving near-perfect accuracy on essays by US eighth-graders (Patterns, 2023; open preprint).

The mechanism is not mysterious. Detectors lean on perplexity, a measure of how predictable text is. Writing by someone working in a second language tends to use a narrower vocabulary and more conventional constructions, which reads as low perplexity, which reads as machine-generated. As the senior author put it, the design of many detectors inherently discriminates against non-native authors. Journalists have since documented international students being accused on that basis (The Markup).

For an international programme this is disqualifying. A detector deployed on a global applicant pool would systematically reject the candidates with the least access, while passing the fluent applicant who used a model well. It automates the exact bias the programme exists to correct.

Some institutions have drawn the sensible conclusion, which is to govern use rather than police style. The NIH prohibits generative AI in the peer review process outright, on confidentiality grounds, and states that it will not treat applications substantially developed by AI as the original ideas of the applicant (NOT-OD-23-149). That is a policy about provenance and originality. It is not a claim that reviewers can tell by reading.

#Separate the proxy from the thing

If the essay was a proxy, the honest response is to ask what it was a proxy for and then measure that directly.

For our programme the answer was narrow and specific. We are not selecting for scholarly promise or writing ability. We are selecting for who will act once their blockers are removed. The programme runs ninety days, carries stipends and data support and mentorship, and is measured on one outcome: verified income earned from a trained skill. Under that definition the useful question is behavioural, not literary.

That reframing matters more than any technique that follows. Most selection debates are really disagreements about the construct being measured, conducted as if they were disagreements about instruments.

#Situational assessment instead of a statement

The replacement is a situational instrument rather than an expressive one. Applicants respond to concrete scenarios drawn from the work itself, of the form: you have done this for thirty days, nothing has worked, what do you do next. There is no correct-sounding answer to reproduce, because the discriminating information is in the choice made, not in how well it is expressed.

This is not novel. It is a well-evidenced field that selection science has used for decades, and medicine has industrialised. Situational judgement tests are embedded in NHS specialty recruitment (NHS England) and in the Scientist Training Programme (NHS NSHCS).

The validity evidence is real, and worth reading before adopting the method. A systematic review and meta-analysis in Medical Education examined situational judgement test validity for selection (Webster et al., 2020). A cohort study of UK general practice training reported that the situational judgement component was the best single predictor of later performance, while also receiving lower face validity ratings from candidates, a tension the authors describe as a justice dilemma (BJGP Open, 2024). Patterson and Zibarras set out the design principles in AMEE Guide No. 100.

Two findings from that literature shaped what we built. Situational instruments and academic measures are complementary rather than substitutable, so they should be scored separately instead of blended into one number. And candidates trust them less than they trust interviews, which means the process has to be explained rather than merely administered.

Responses are scored into a composite readiness signal with risk flags and a clear recommendation to admit, waitlist or reject. That recommendation goes to a human. It is not a decision.

#Capability is not one number

The second change was to stop treating capability as a single quantity.

We assess three dimensions separately, because each one alone produces a failure we had watched happen repeatedly. Strong technical ability without commercial awareness produces someone who can build but cannot earn. Commercial instinct without technical depth produces someone with nothing to sell. Both without interpersonal skill produces someone nobody wants to work with twice. Averaging the three hides exactly the imbalance that predicts failure, so imbalance is surfaced explicitly and progression requires a threshold on each.

This is where the essay was most misleading. A well-written statement flatters the first dimension and says nothing reliable about the other two.

#Make the outcome the proof

The final change is the one most programmes will resist. We do not treat completion as success. Graduation requires verified income earned from the trained skill, paid by an external party, before day ninety. Not a mock client, not an internal payment, not an unrelated hustle.

That single rule does more for selection integrity than any assessment redesign, because it closes the loop. If the intake instrument is admitting the wrong people, the outcome data says so within a quarter. An essay-based process can be wrong for years without ever finding out.

#What the model is and is not allowed to do

The scoring uses a language model, which introduces its own integrity problem. We wrote the limits down before shipping.

The model classifies, flags risk, personalises the work assigned, and tunes difficulty. It is expressly forbidden from overriding programme rules, granting exceptions, or determining who graduates. Verified income is the only proof of outcome. If the model layer is unavailable, scoring falls back to deterministic rules and logs the downgrade, so an intake cohort is never blocked because an external service is down.

The distinction we hold to is that the model is a referee and never a judge. A system that can state plainly what its model may not decide is one you can defend to a regulator, a funder, or the applicant it just turned down. That last audience is the one most teams forget, and the one most likely to ask.

#What this costs

An honest accounting, because the alternative is not free.

Scenario design is expensive and perishable. Each scenario needs to be written from real situations, calibrated, and retired once it circulates. Essays scale to any volume at no marginal cost; a scenario bank does not.

Scoring is harder to explain. A rejected applicant can be shown their essay. Explaining a composite readiness signal takes work, and if you cannot explain it you should not deploy it.

Candidates trust it less, as the general practice cohort study found. That is a communication obligation, not a reason to avoid the method.

And the evidence here is one programme, not a study. The selection science I cite is peer-reviewed; our implementation of it is not. Treat the method as transferable and our numbers as anecdote.

#When not to do this

Do not adopt this if the written statement is measuring something you genuinely need. A creative writing fellowship should assess writing. A policy programme that will require drafting under time pressure should test drafting.

Do not adopt it if you cannot define your outcome. The reason we could replace the essay is that we have an unambiguous, externally verifiable measure of success. A programme whose outcome is "increased capacity" has nothing to calibrate an instrument against, and will build something arbitrary and call it rigorous.

Do not adopt it at small volumes. If you receive forty applications, interview all forty. Structured human judgement beats any instrument at that scale.

And do not bolt a detector on instead. That is the one option the evidence rules out.

#Key takeaways

  • The written statement was a proxy for clarity, seriousness and language command. Generative tools broke the correlation, so it now chiefly measures access.
  • Detection cannot restore it. Detectors misclassified over half of non-native TOEFL essays as machine-written while scoring near-perfectly on US eighth-grade work, so a detector deployed globally automates the bias the programme exists to correct.
  • Identify what the essay was standing in for and measure that directly. For a programme selecting for action, the useful instrument is situational rather than expressive.
  • Score capability on separate dimensions rather than one blended number. Averaging conceals the imbalance that predicts failure.
  • Tie graduation to an externally verified outcome. It is the only mechanism that tells you whether your intake instrument is working.
  • Write down what the model may never decide, before shipping. Then you can defend the process to whoever asks, including the person it rejected.

#References

  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7). doi:10.1016/j.patter.2023.100779. DOI · open preprint · plain-language summary
  • AI Detection Tools Falsely Accuse International Students of Cheating. The Markup, 2023. Link
  • National Institutes of Health. The Use of Generative Artificial Intelligence Technologies is Prohibited for the NIH Peer Review Process, NOT-OD-23-149. Link · NIH Office of Science Policy on AI
  • Kaplan. Colleges Increasingly Clarify Rules on GenAI Use in Admissions Essays, But Majority Still Keep Applicants Guessing, 2025. Link
  • Pew Research Center. About a quarter of US teens have used ChatGPT for schoolwork, double the share in 2023, 2025. Link
  • Webster, E. S., Paton, L. W., Crampton, P. E. S., Tiffin, P. A. (2020). Situational judgement test validity for selection: A systematic review and meta-analysis. Medical Education, 54(10). doi:10.1111/medu.14201. DOI
  • New evidence on the validity of the selection methods for recruitment to general practice training: a cohort study. BJGP Open, 2024. Link
  • Patterson, F., Zibarras, L. Situational judgement tests in medical education and training: Research, theory and practice. AMEE Guide No. 100. Link
  • NHS England. Situational Judgement Test, medical specialty recruitment. Link · Scientist Training Programme longlisting and SJT. Link
  • Boston Globe. Students are using AI to write scholarship essays. Does it work?, 2025. Link
  • Hechinger Report. Students try using AI to write scholarship essays, with little luck. Link

Shipping something like this?

Everything above came out of building and running the thing, not researching it. If you are hitting the same problems, I work with product and engineering teams on exactly this, architecture reviews, technical strategy, and staying on through implementation.