Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Netflix’s E2EGraph is a proposed way to connect user experience, client applications, services, and infrastructure in one shared model. In a QCon London 2026 presentation, Netflix engineers Prasanna Vijayanathan and Renzo Sanchez-Silva described an ontology-driven graph intended to help engineers trace a user-visible regression across system boundaries. The talk also outlined AutoSRE and automated mitigation as future directions—not as proven production capabilities.

Why end-to-end observability is a correlation problem

A viewer may experience a slow or degraded feature even when no single service looks obviously unhealthy. The cause could lie in a client release, a network component, a backend dependency, an experiment, or an interaction among them. At Netflix’s scale, telemetry comes from hundreds of client platforms, microservices, and infrastructure components, while metrics, traces, and logs are divided across systems and teams.

That division makes a question such as “Is this user-visible regression caused by the client, the network, or a backend dependency?” difficult to answer. Engineers need to connect evidence from different places and establish that the records refer to related parts of the same user journey. The presentation frames E2EGraph as a response to this integration and meaning problem, rather than simply as another place to store telemetry.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What E2EGraph is designed to represent

The session abstract describes a graph in which entities are nodes and their relationships are edges. The model spans a user experience from session and client through services and network components, with operational context attached to relevant nodes or connections.

Graph element Examples described in the presentation What it helps connect
Nodes User sessions, client applications, microservices, network components The entities participating in a user journey or operational event
Edges Requests, user interactions, dependencies How one entity calls, depends on, or interacts with another
Attributes Latency, error rate, quality of experience (QoE), version, geography How an entity or relationship behaved in a particular context

The described inputs include client telemetry, server logs, traces, infrastructure metrics, experiments, and deployments. A graph representation can make relationships explicit: a session uses an application, the application makes an API call, and that call depends on a service. The aim is to let analysis follow those links instead of relying on engineers to manually reconcile identifiers and dashboards.

Why an ontology matters

A graph alone does not guarantee that separate systems mean the same thing by a term or represent a relationship consistently. E2EGraph’s ontology is described as a shared semantic layer: a formal specification of entity types, properties, and relationships that can normalize different data sources and support consistent reasoning across them.

For example, the talk identifies concepts such as a user session, API call, deployment event, experimentation, and QoE regression. InfoQ’s March 18, 2026 event report explained graph facts as subject-predicate-object triples. Its examples associate an API gateway with an application type and an owner, and connect an incident to an affected gateway. In such a model, systems can represent not just that two records exist, but what each record is and how the records relate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

InfoQ also reported that the presentation described 12 operational namespaces, with examples including Slack, alerts, metrics, logs, incidents, E2E, and Harvest. Those are conference-reported presentation details; they do not establish a separately verified production schema. The important design distinction is that the ontology is meant to give diverse data a common vocabulary, while the graph records entities and their connections using that vocabulary.

How the graph could support regression investigation

The abstract describes graph snapshots around regression events so engineers can compare healthy and degraded states over time. In principle, an investigation could start with a QoE change, traverse related sessions, client versions, requests, dependencies, experiments, and deployments, and then compare the connected context with a healthier period.

  1. Start from an observed regression. Identify the affected user experience and the relevant time window.
  2. Follow modeled relationships. Trace from sessions and client applications through requests and dependencies toward services and infrastructure.
  3. Use context attached to the graph. Examine attributes such as latency, error rate, client version, geography, and QoE alongside events such as deployments or experiments.
  4. Compare snapshots. Look for meaningful changes in the connected graph between healthy and degraded states, rather than treating each telemetry source in isolation.

This is the investigation logic implied by the presentation, not a published step-by-step operating procedure or a claim that every listed data source is already connected in production. The abstract’s example query—“Why is TV UI lolomo TTR regressing in the latest version?”—illustrates the kind of user-experience question the model is meant to address.

AutoSRE is described as a layer above the graph

The presentation positions AutoSRE as a planned analysis layer that would use the graph to organize operational investigations. A coordinator agent would break a question into tasks, specialized agents would query graph-connected domains such as metrics, alerts, experiments, client platforms, events, and deployments, and the coordinator would synthesize the evidence into likely root causes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a TV interface regression, the abstract gives possible explanations including a client rollout, a misconfigured experiment, or a backend dependency regression. In this design, the graph provides relationships and context for the agents to inspect; the proposed coordinator assembles findings across domains. This is a described architecture, not evidence of measured autonomous diagnosis in production.

What the incident figures do—and do not—show

InfoQ’s March 18, 2026 report said a recent Netflix incident took four hours from the initial alert to resolution. It also reported that nine teams and more than 30 engineers worked on that incident and three related incidents. Those figures are attributable to InfoQ’s event report; the available source set does not include an independently linked Netflix incident post verifying them.

The figures illustrate the coordination burden the presentation’s approach seeks to reduce, but they do not demonstrate that E2EGraph or AutoSRE shortened that incident. The report does not establish that the graph was used in the incident, nor does it provide a before-and-after comparison.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Prediction and remediation remain roadmap intentions

The talk’s stated roadmap includes predicting risk from patterns in failing subgraphs, propagation paths, and risky combinations of versions and experiments. It also describes possible mitigations such as targeted rollback, traffic shifting, feature-flag changes, and capacity adjustment. InfoQ reported broader future aims of automating root-cause analysis, enabling auto-remediation, and building self-healing infrastructure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These are intended directions, not demonstrated outcomes. The presentation materials described in the available sources do not establish that these mitigations run automatically, how approval would work, or what safeguards would prevent an incorrect diagnosis from triggering a harmful change. Those distinctions matter: a system that surfaces graph-backed evidence for an engineer is materially different from one authorized to change production traffic or configuration on its own.

What is not established publicly

QCon’s session page says slides are unavailable. The available public material does not provide an implementation architecture, graph size, query latency, diagnosis accuracy, evaluation method, adoption level, or measured operational outcomes. It therefore supports an explanation of the proposed model and direction, but not claims about its production maturity or effectiveness.

For teams evaluating a similar design, useful questions include whether telemetry is merely correlated or normalized under explicit domain semantics; whether analysis spans the full client-to-infrastructure path; whether healthy and degraded states can be compared over time; whether root-cause explanations expose inspectable graph evidence; and whether remediation is advisory, human-approved, or automated. These are evaluation criteria suggested by the presentation’s goals, not benchmark results or comparisons reported by Netflix.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.