Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universally accepted score that tells regulators an AI model is “safe” on one side and “dangerous” on the other. In the European Union, the clearest numeric signal is a legal presumption: a general-purpose AI model trained using more than 1025 floating-point operations (FLOP) is presumed to have high-impact capabilities. That threshold prompts closer regulatory scrutiny; it is not proof that a model will cause harm.

What the EU’s 1025-FLOP threshold means

Under Article 51 of the EU AI Act, a general-purpose AI (GPAI) model is classified as presenting systemic risk if it has high-impact capabilities, assessed with appropriate technical tools and methods, or if the European Commission determines that it has equivalent capabilities or impact under Annex XIII. The Act presumes high-impact capabilities when cumulative training compute is greater than 1025 FLOP. Read Article 51 and Annex XIII of Regulation (EU) 2024/1689.

Compute is a useful early-warning measure because it can be quantified. But it is a proxy, not a direct measure of what a model can do or the harm it will cause. The Commission can update thresholds and benchmarks as technology changes. A provider whose model crosses the threshold must notify the Commission and may give reasons it should not be classified as systemic risk; the Commission assesses those reasons. The alternative designation route also means a model below the compute threshold is not automatically outside scrutiny.

A separate figure in Commission guidance is an indicative 1023-FLOP training-compute criterion for identifying certain GPAI models. It concerns GPAI identification, not the 1025-FLOP presumption for systemic-risk capabilities, and it is not an absolute rule: the Commission says generality and capability matter and exceptions are possible. See the European Commission’s GPAI provider guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Two routes to a systemic-risk designation

Route What is assessed How the threshold or decision is set What the route can miss or capture
Compute presumption Cumulative computation used to train the model, measured in FLOP. The AI Act sets the presumption at more than 1025 FLOP; the Commission can update thresholds. It gives regulators a measurable trigger, but compute need not track capability perfectly as algorithms and hardware improve. Providers may present reasons against classification, which the Commission evaluates. AI Act, Article 51; Commission guidance.
Capability or impact designation Technical evidence about capabilities and the factors in Annex XIII, including model size, training-data quality or size, modality, and market reach. The Commission may determine that a model has equivalent capabilities or impact. A 2025 research proposal suggests benchmark-based measurement, but it is not an adopted legal scoring system. It can reach models below the compute threshold or account for broader impact, but depends on technical evaluation, benchmark choices, reference thresholds, and regulatory judgment. AI Act, Annex XIII; EU Publications Office proposal.

How researchers propose measuring capability

A report published by the EU Publications Office on 8 October 2025 proposes combining results from a diverse set of benchmarks into a composite measure. Examples include MMLU-Pro, GPQA-diamond, MATH-level-5, and HumanEval. The proposal derives benchmark weights using principal component analysis (PCA), then has the enforcement authority set a threshold relative to a reference model, informed by legal, policy, and risk considerations. It recommends expert oversight of benchmark selection and updating the method every six months. Read the report.

This is a proposed way for authorities to operationalize capability, not a settled EU score or universal danger meter. Benchmark performance can show that a model handles a particular class of tasks under test conditions. It does not, by itself, establish the probability of real-world harm: that also depends on access, deployment, safeguards, users, and scale. The cited proposal does not establish that a composite score reliably predicts harm across every deployment.

Why reach and deployment matter too

Systemic risk is not limited to a model’s training compute or frontier capability. The European Commission Joint Research Centre’s reach study explains that widely used models can shape users’ information environment, including through bias-related effects, even when they are not at the technological frontier. It discusses measuring direct interaction through user interfaces and APIs and proposes user-count and reporting measures to complement capability, safety benchmarks, and compute. These reach measures are proposals in the report, not a universal adopted threshold. Read the Joint Research Centre report on GPAI model reach.

The AI Act also includes a reach-related presumption: Annex XIII identifies 10,000 registered business users in the Union as relevant to high impact on the internal market. That is one factor in the Act’s assessment, not a substitute for the separate compute presumption. The Commission’s guidance and the Act text are the sources for the applicable legal interpretation. AI Act, Annex XIII.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The European Systemic Risk Board’s 2026 warning on systemic cyber risks from frontier AI models illustrates why regulators consider potential effects on sectors and connected systems, not just an abstract model score. Read the ESRB warning.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What systemic-risk classification requires

For providers of GPAI models classified as presenting systemic risk, the Commission’s guidance summarizes obligations that include:

  • Evaluating the model with standardized protocols and state-of-the-art tools.
  • Conducting and documenting adversarial testing.
  • Assessing and mitigating systemic risks.
  • Tracking and reporting serious incidents and corrective measures.
  • Providing adequate cybersecurity for the model and its physical infrastructure.

The Commission states that GPAI obligations began applying on 2 August 2025. Its guidance says full compliance enforcement, including fines, begins on 2 August 2026; models already placed on the market before 2 August 2025 have until 2 August 2027 to comply. Check the Commission’s guidance for the current timing and obligations.

Is 1025 FLOP a global definition of dangerous AI?

No. It is an EU legal presumption for GPAI models, not an international consensus or a finding that every model above the line is dangerous. The sources cited here establish the EU framework and research proposals around it; they do not establish a comparable threshold scheme for the United States, United Kingdom, or other jurisdictions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.