iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Whitehall’s recruitment system can identify CV keywords and interview performance without proving that a technologist can do the job. Laura Gilbert, the former head of AI for government, told MPs that practical, job-relevant testing is needed—starting with a four-hour coding exercise based on a realistic scenario.
What Laura Gilbert told MPs
Giving oral evidence to the House of Commons Science, Innovation and Technology Committee on 25 March 2025, Gilbert said civil-service hiring for technical roles is “really hit and miss”. Her criticism was about validity: a candidate may list the right programming languages and terminology yet still be unable to deliver the work required.
“It’s very difficult to hire technologists well. The way the civil service hires for this sort of role is not suitable for this sort of purpose. It’s really hit and miss,” Gilbert said.
PerformanceWindows Errors? Fix Them Before They SpreadDriversCrashes, No Sound, or Screen Glitches?PerformancePC Slower Than It Used to Be?Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
She said the system does not reliably assure that applicants can perform the job. In her view, continuing to use behaviour-based interviews designed for general civil-service recruitment on technical posts leaves a basic capability question unanswered.
#1 Best Overall
“The system doesn’t have a way to hire that assures that people [can] do the job. Until that changes, it’s going to continue to be very difficult,” Gilbert said.
The alternative: test the work, not just the CV
A four-hour, realistic coding exercise
Gilbert contrasted standard screening with the process used for applicants to her AI incubator. Candidates completed a four-hour coding test built around a real scenario. The point was not to reward familiarity with recruitment jargon; it was to observe how someone approaches a task resembling the work they would actually undertake.
This is evidence of a different selection principle, not proof that one test is a complete hiring system. A practical exercise can reveal coding and problem-solving ability that a CV or generic interview may miss, but it must be designed, marked and supported consistently.
Free tools Windows power users keep installed
One-click scans. No signup required.
How the approaches differ
| Question | Typical CV and behaviour-interview approach | Gilbert’s practical-testing approach |
|---|---|---|
| Evidence of job validity | Keywords, claimed experience and answers to competency questions | Observed performance on a realistic coding scenario |
| Candidate burden and speed | Usually spread across application and interview stages; exact burden varies | A defined four-hour exercise for AI-incubator applicants |
| Consistency between departments | Can vary by department, panel and interpretation of behaviours | Could be standardised if the task, scoring and safeguards were shared |
| Roles covered | Can be used for many jobs, but may not show technical ability well | Strongest fit for hands-on coding; separate methods are needed for leadership, data, security and non-coding work |
| What it cannot establish alone | Technical delivery is often inferred rather than demonstrated | It does not by itself measure long-term leadership, collaboration, retention or performance on government legacy systems |
Why this matters to the Government Digital Service
The complaint arrived as ministers were planning a larger reorganisation of government technology. A January 2025 blueprint set out a six-point reform plan and a refreshed Government Digital Service (GDS) intended to bring existing digital, data, artificial-intelligence and related teams into a more unified capability.
That ambition increases the cost of weak selection. If posts are labelled “digital” but employees do not possess the capability those labels imply, expanding the number of digital staff will not produce the expected delivery capacity. Recruitment is therefore a foundation for the refreshed GDS, rather than a side issue.
The available evidence does not describe a final GDS hiring framework or say that every technical applicant will sit the same coding test. It supports a narrower conclusion: hiring methods need credible, role-relevant evidence of competence, with practical tests as one possible component.
The £45 billion figure—and what it does not mean
The government’s State of digital government review, published by the Department for Science, Innovation and Technology and GDS in 2025, estimated £45 billion per year in unrealised savings and productivity benefits from full digitisation. That is a government estimate of potential benefits, not a guaranteed cash saving or an amount already available to spend.
Recommended Free Tools
- Public-sector digital technology spending is reported at more than £26 billion a year.
- The public sector has nearly 100,000 digital and data professionals.
- Digital technology accounts for roughly 4–7% of public-sector spending, according to the same 2025 review.
Those figures show why capability decisions matter financially, but they do not demonstrate that changing recruitment alone will unlock £45 billion. The estimate depends on successful delivery across departments, not simply on hiring more programmers.
Best Value
Recruitment is one part of a larger delivery problem
Legacy technology
Richard Pope, who also gave evidence to the committee, pointed to the government’s stock of unmaintained technology: “There’s a lot of unmaintained technology the government relies on to do its job. It’s not good enough at the moment, and we obviously need to fix that.” New hires must be able to work safely around ageing systems while departments fund remediation and replacement.
Data and interoperability
Digital services cannot deliver joined-up outcomes when teams lack clear rules for exchanging data, ownership of interfaces and reliable information. A coding exercise may show implementation skill, but cross-government data governance still requires organisational decisions.
Leadership, pay and retention
Hiring a capable engineer is not the same as keeping one. Government must compete on pay and career progression, provide credible technical leadership and make security-clearance processes workable. No public statistic gives a measured count of government technologists who are unsuitable, and no controlled evaluation proves that coding tests improve hiring outcomes.
Different roles need different evidence
A coding task is naturally relevant to software engineering and some AI roles. It is not a universal test for senior technology leadership, product management, architecture, data governance, user research, cyber security or commercial roles. A coherent model would match assessment to the work: practical technical tasks where appropriate, structured job simulations and evidence-based interviews for other responsibilities.
What a credible reform would need to answer
- Define the capability. Each advert should specify the technical outcomes and level expected, rather than relying on broad “digital” wording.
- Use job-relevant exercises. Tasks should resemble real work, allow reasonable accessibility adjustments and be scored against published criteria.
- Standardise without flattening roles. Shared principles can improve consistency across departments while allowing different tests for coding, data, leadership and non-coding posts.
- Protect fairness and security. Panels need training, conflict-of-interest controls and processes that account for reasonable adjustments and clearance requirements.
- Measure the result. Departments should track candidate experience, time to hire, retention and on-the-job performance rather than assuming that a new test works.
- Connect hiring to delivery. Workforce plans must sit alongside legacy-system remediation, data-exchange rules, technical leadership and funding.
Bottom line for applicants and taxpayers
Gilbert’s testimony identifies a specific weakness: CV screening and behaviour-based interviews do not reliably demonstrate that a technical candidate can perform technical work. A four-hour, realistic coding test is a concrete alternative for suitable roles, but it is not a standalone solution. The refreshed GDS and the government’s projected digital benefits will depend on a broader system that hires, supports and retains the right people while fixing legacy technology and cross-government delivery.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

