iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
AI providers should not be the sole judges of whether their own systems are trustworthy or safe, especially when a failure would be costly. Their internal evaluations still matter, and in practice they are often the most detailed tests available. The useful question is not whether a company may test its own model, but when an independent check should be added on top, and how to design it so it adds real scrutiny.
Can AI companies be trusted to test their own models?
Usually they are the best-placed party to run the first round of testing, and that is not a criticism. A developer knows how the model was trained, what data went into it, what safeguards were built in, and where the known weak spots are. The problem is structural. The same organization that built a system and benefits from its commercial success is also the one deciding whether the results look good enough to release, describe publicly, or sell. Even well-intentioned teams can miss failures they are not motivated to find.
So the answer is that provider testing can be trusted as one input, but it is weaker as the only input. The stronger the consequences of an error, the more a reader should expect an outside check.
What the main frameworks say
Three reference points are useful here, and they point in the same direction without all imposing the same obligations.
#1 Best Overall
NIST AI Risk Management Framework: voluntary guidance
The U.S. National Institute of Standards and Technology published its AI Risk Management Framework (AI RMF 1.0) in January 2023. NIST describes the framework as voluntary. It offers guidance for incorporating trustworthiness considerations into how AI is designed, developed, used, and evaluated. It does not create a legal duty to hire third-party auditors.
The framework does, however, make a direct point about independence. In AI RMF 1.0 (2023), NIST states: “Processes for independent review can improve the effectiveness of testing and can mitigate internal biases and potential conflicts of interest.” That sentence is the clearest official statement that internal testing has a built-in limitation, and that outside review is a recognized way to address it.
NTIA’s accountability report: internal testing is maturing, and both approaches have a role
In its Artificial Intelligence Accountability Policy Report (March 2024), the U.S. National Telecommunications and Information Administration (NTIA) makes a balanced case. It reports that internal evaluations benefit from access to relevant material, and that they are currently more mature and robust than independent evaluations. That is a meaningful admission in favor of provider testing: the people with the most access and the most experience with the system are often the ones best able to test it.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #2
The same report notes calls for independent evaluations where warranted, as a check on false claims and on risky AI. It presents the two approaches as potentially complementary rather than as substitutes. The policy takeaway is not that companies are untrustworthy, but that independent evaluation has to mature alongside internal practice before it can carry the weight placed on it.
EU AI Act: specific duties for systemic-risk general-purpose models
The EU AI Act sets more specific duties, but only for a defined group. Article 55 applies to providers of general-purpose AI models with systemic risk. According to the text shown by the European Commission’s AI Act Service Desk (consolidated text as of 27 July 2026), those providers must, among other things:
- perform model evaluation using state-of-the-art standardized protocols and tools;
- document adversarial testing;
- assess and mitigate systemic risks;
- report serious incidents; and
- maintain cybersecurity protections.
Recital 114 of the same Act says that the necessary model evaluations may use internal or independent external testing. It does not require independent external testing in every case. Readers should not assume that the EU Act always requires an outside auditor. What it requires is a documented, rigorous evaluation process, and it leaves room for how that evaluation is staffed. Legal obligations also depend on a provider’s role, how a model is classified, and the jurisdiction, so anyone relying on this should check the current consolidated text.
Rank #3
Why internal testing still carries weight
Dismissing provider testing would make the problem worse, not better. Internal teams can run tests that outsiders cannot, because they have:
- access to training data documentation, model internals, and the deployment context;
- the ability to run repeated, targeted tests during development rather than only at release;
- direct channels to fix problems they find, which shortens the time between discovery and remediation.
Those advantages are why NTIA describes internal evaluation as more mature today. An external reviewer who sees only a finished model through an API has a narrower view, and an independent check that lacks system context may test the wrong things. Independence improves credibility; it does not automatically improve technical depth.
Internal and independent evaluation compared
The table below compares the two approaches on the axes that matter for trust. It draws on the distinction NIST and NTIA make between internal access and maturity on one hand, and the conflict-mitigation role of independent review on the other. It is a reasoning aid, not a measured scorecard.
| Axis | Internal (provider) evaluation | Independent evaluation |
|---|---|---|
| Access to development data and system context | Strong: direct access to training, configuration, and internal testing records | Limited unless the provider grants access; often limited to the released system |
| Independence from commercial incentives | Weakest: the evaluator and the product owner share the same incentives | Stronger by design, though independence depends on the reviewer’s own funding and arrangements |
| Expertise and maturity | Described by NTIA (March 2024) as currently more mature and robust | Described by NTIA (March 2024) as less mature at present; capacity varies |
| Reproducibility and transparency | Depends on what the provider discloses; not stated in the cited sources as a fixed standard | Depends on the reviewer’s documented methods and whether results can be replicated; not stated in the cited sources as a fixed standard |
| Ability for outsiders to validate claims | Low unless methods and results are published | Higher where reviewers can publish findings or the provider shares access |
| Cost, timeliness, and scope | Usually faster and more iterative; cost figures not stated in the cited sources | Usually slower to arrange and narrower in scope; cost figures not stated in the cited sources |
Who checks the safety claims, and when independent scrutiny is warranted
The better question is not whether to have an outside check, but how much independence a given system needs. Use the consequences of error as the main guide. An independent review is most justified when several of the following conditions apply:
- High-impact use. The system informs decisions about health, finance, employment, education, legal rights, or public services.
- Serious or hard-to-reverse harm. A failure could affect many people or cause damage that cannot easily be corrected.
- Conflicts of interest. The provider has strong commercial reasons to show that a system is safe, or a release timeline pressures the evaluation.
- Public safety or capability claims. The provider publicly describes a system as safe, unbiased, or reliable, and readers cannot check those statements themselves.
- Frontier or systemic-scale capability. The model is general-purpose and widely deployed, the kind of system the EU AI Act’s Article 55 targets.
- Limited provider visibility. The risks involve third-party deployments or misuse that the developer may not see directly.
When few of these apply, a well-documented internal evaluation may be proportionate. When several apply, internal testing alone is unlikely to be enough, and an independent check should be expected.
How to combine provider knowledge with independent scrutiny
Independent review adds the most value when it is designed to fix the weaknesses of internal testing rather than repeat it. A practical arrangement has four parts:
Best Value
- Set the scope before testing starts. Define which capabilities, risks, and user groups are covered, and what would count as a failure. Write this down before the results are known.
- Give the reviewer defined access. Agree in advance what the reviewer can see, such as documentation, evaluation logs, or query access, and state any limits in the published report.
- Publish methods and limits. Report what was tested, what was not, and how reviewers were selected and funded. Credible review depends on reviewers being able to show their work.
- Track findings to resolution. Record each material finding, the fix, and whether the fix was checked. Unresolved findings should be visible to the public.
Each step addresses a specific failure mode of self-grading: unclear scope lets results be framed favorably, limited access hides problems, unpublished methods cannot be checked, and unresolved findings disappear from view.
What this argument does not claim
The case here is not that every AI system needs an outside audit, that provider evaluations are worthless, or that voluntary guidance is a legal requirement. NIST’s framework is voluntary. The EU AI Act’s evaluation duties apply to a defined category of providers, and its recitals allow internal testing. NTIA itself treats internal evaluation as currently the more mature approach. The defensible position is narrower: when the consequences of error are high, or the provider has a conflict, an independent check should be part of how trust is established, and provider testing should be designed so that outsiders can verify it.
Readers who want to follow this topic should watch how NIST’s AI Resource Center develops operational guidance for testing, evaluation, verification, and validation, and check the current consolidated EU AI Act text for any system they are assessing.
“,
“faq”: [],
“bottom_line_html”: “”,
“seo_title”: “Why AI Companies Shouldn’t Be the Only Ones Grading Their Own Systems”,
“meta_description”: “AI providers’ own tests matter, but independent review reduces conflicts of interest. What NIST, NTIA and the EU AI Act say, and when outside checks are warranted.”,
“excerpt”: “AI providers’ internal testing has real advantages, but NIST and NTIA both point to independent review as a way to reduce bias and conflicts of interest, especially for high-stakes systems.”,
“tags”: [“AI governance”, “AI safety”, “AI evaluation”, “NIST AI RMF”, “EU AI Act”, “AI accountability”],
“research_used”: true}
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

