Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Sentinel is a custom pattern for answering one morning question in one message: is the data behind our dashboards current, complete, and correct? The author, writing in a first-person DEV Community post dated September 29, 2026 under the handle gentjan_likaj, describes collecting health signals from every layer of an AWS-centered pipeline, storing them as JSON in Amazon S3, and having a short-lived AI agent turn them into a single Slack digest. The central claim is that a green orchestration run does not establish that delivered data is complete or correct.

This is one practitioner’s account of what they built and what they say it changed. It is not an independent evaluation, and the post publishes no measured accuracy, alert-noise, or time-saved figures. The sections below explain how the pattern works, which design decisions carry the most weight, and what remains unproven.

The question that starts the morning

The trigger is a stakeholder message that every data team recognises: “The dashboard looks off. Is the data updated?” In the author’s description, answering it meant opening Airflow, checking AWS Glue job runs, looking at dbt results, and then opening Tableau to see whether an extract had refreshed. Each tool could report its own success, yet nobody had a single view of whether the numbers a business user was about to read were trustworthy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sentinel’s goal is to move that check forward in time. Instead of an analyst reconstructing the state of the stack after a complaint, a digest arrives each morning and states either that everything is clean or which problems exist, who owns them, and what probably caused them.

Why separate tools hide the answer

The example pipeline takes data from APIs and databases, lands it through AWS Glue into Redshift, transforms it with dbt, writes a further Redshift layer, and feeds Tableau. Each layer has its own vocabulary for success, and the author stresses that a pipeline can look healthy while the output is wrong. A job can finish without error and still deliver fewer rows than usual, a KPI can move in a way that signals a broken join, and a report can show values that no longer match the source of truth.

Layer What the tool reports as healthy What a healthy status does not prove
Airflow The DAG or task run completed The run finished late, needed retries, or breached a time-of-day SLA
AWS Glue The job run succeeded Which partition or file was processed, and whether row volume is in line with recent weeks
dbt Models built without errors Sources are fresh, and the output has the expected volume and values
Redshift Tables loaded and queries returned The loaded data matches the business metrics it is supposed to represent
Tableau Extract refresh or view loaded The figures on the dashboard agree with the benchmark or source of truth

The point of the table is not that these tools are unreliable. Each one is answering a narrower question than the stakeholder is asking.

What Sentinel collects

The author groups the signals by where they originate. Each group below is a separate collector, and each writes its own health record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Orchestration: Airflow

The Airflow collector records the latest pipeline state, the duration of the run relative to its average, the number of task retries, the owner, and SLA status. SLAs are defined as time-of-day deadlines, so a run that eventually succeeds can still be flagged if it finished after the deadline. The author also describes a failure callback that writes a meaningful error line to S3 on a best-effort basis. Best-effort means the error text is useful when it arrives, but the design does not guarantee that every failure produces one.

Ingestion: AWS Glue

For each Glue job, the collector captures recent run status, duration, and the error message. Where a job runs many times a day, the author retains per-run parameters so the digest can say which partition or file a failed run was processing, rather than only that “the job failed.”

Transformation and sources: dbt

The dbt collector records model executions and errors. It also checks source freshness and compares row volumes against the same weekday in the prior week. Comparing against the same weekday, instead of yesterday, is a deliberate choice that accounts for normal weekly cycles in business data, although the post does not say how it handles holidays or other irregular days.

Business metrics

Core KPIs such as costs, leads, sessions, and orders are compared with the same day in the previous week. The author says very small values are skipped, because a small absolute movement can look dramatic in percentage terms without meaning anything. The post does not supply a threshold formula, so readers should treat the cut-off for “small” as a decision each team must make and document.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reconciliation: the North Star KPI

A North Star check compares a report with a benchmark. This is the signal aimed at drift from a source of truth or at historical restatements, where a report still loads and its totals have quietly changed. It is the check most likely to surface problems that no job status would ever show.

Consumption: Tableau

The Tableau collector identifies failed extract refreshes and the datasource owner, so a failure can be routed to a person rather than posted to a general channel.

How the health data travels

Each collector writes JSON health metadata to S3, and every consumer reads from that same store. The author presents this as a way to decouple producers from consumers. Payloads stay inspectable, they can be replayed, and any HTTP-capable consumer can use the same data. These are design benefits the author claims, not properties tested by anyone else.

A small internal HTTP API sits in front of the bucket. It reads a requested health file from S3 and returns it as JSON. Keeping that layer thin means the digest agent never needs direct S3 credentials to reason about the stack, and other reports or agents can call the same endpoints. The author notes that this is also what allows the health metadata to feed other reports, not only the morning message.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the morning digest is produced

After the collectors finish, Airflow starts a short-lived agent session with a prompt kept in version control and shell access to the gateway. The sequence the author describes is:

  1. Call the gateway endpoints for each collector’s health file.
  2. Filter the records for failures and anomalies, using deterministic code rather than the model.
  3. Map each remaining issue to its owner, using the owner fields captured by the collectors.
  4. Reduce noisy error text to a probable root cause, so several symptoms from one upstream failure appear as one item.
  5. Compose one Slack post. If nothing qualifies, the post is an all-clear. Otherwise it is a grouped list of issues with owners attached.

The output format matters for triage. The author’s stated goal is that the first thing a data engineer reads each morning is one message with a clear state, rather than a sequence of tool tabs.

Operating lessons for putting an AI agent in the loop

The most useful part of the post, for teams considering a similar build, is what the author says went wrong or had to be designed around. These lessons apply to any pipeline that uses a language model to summarise operational data.

Filter the full payload before the model sees it

The author’s key safeguard is not to ask the model to decide what to drop from a large input. Every record is filtered deterministically first, and the post names jq as the tool for that step. Only the smaller set of qualifying records goes to the model for summarisation. The reason is concrete: in an earlier iteration, model-side truncation omitted records and produced a false clean report. A digest that says “all clear” when the stack is not clear is worse than no digest, because it removes the reason to look.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Treat empty or broken replies as failures

If the model returns an empty response, or a malformed one, the run should fail visibly. A silent empty reply that is read as “nothing to report” is the same false all-clear problem in a different form.

Do not blindly retry billable, non-idempotent steps

Posting to Slack or calling a billable model endpoint is not safe to repeat automatically, because a retry can produce a duplicate post or a second charge. The author’s recovery path is the next scheduled run, not an immediate retry. The trade-off is that a failed morning digest can leave the team without a message until the next cycle, so the failure itself must be visible.

Tear down agent sessions

Sessions should be closed after success and after failure. A session that is left running costs money, holds access, and can contaminate later runs.

Version-control the prompt

The prompt is code. Changes should go through review like any other change, because a small wording change can alter which issues are reported and which are dropped.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Report partial outages instead of suppressing the digest

When one source is unavailable, the digest should say so, and then deliver the results it does have. Suppressing the whole message when one gateway call fails would hide the healthy parts of the stack and would also hide the fact that the monitor itself is partly blind.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the author reports, and what is not established

The author reports that morning triage now happens in one Slack message. Silent problems such as low row volume and KPI or report drift became visible, owners are tagged on each issue, and the health metadata can feed other reports and agents. These are the author’s own observations after running the system, not measured outcomes.

The post does not establish any of the following:

  • Alert accuracy, including how many flagged issues were real problems.
  • Reduction in noise compared with the previous monitoring approach.
  • Mean time to detection for data incidents.
  • Time saved by the team each morning.
  • How the approach compares with commercial data observability products or other architectures.

The example digest line “~50% of last week” appears in the post as an illustration of how a volume signal might read. It is a sample message, not a measured result from the author’s stack.

Building or evaluating a similar digest

Because the post does not compare alternatives, the most useful way to read it is as a set of requirements. A team assessing its own version, whether built in-house or bought as a product, can check each of these:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Does it cover every layer a stakeholder question crosses, including ingestion, transformation, and reporting, not only job status?
  • Does it check freshness, row volume against a comparable period, business KPIs, and report-level values against a benchmark?
  • Does each issue carry a named owner and enough error context to act on?
  • Is filtering deterministic and auditable, with the model working only on the qualifying set?
  • What happens when a source is down? Does the digest say so?
  • Can a failed delivery be detected, and are retries of non-idempotent steps prevented?
  • What is the operational overhead of the collectors, the gateway, the session lifecycle, and prompt maintenance?

A team that can answer all seven has a credible design. A team that can answer only the first two has a status page with a chatbot in front of it.

The post’s closing formulation is the one most worth keeping in mind: “Green pipelines don’t mean correct data.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.