Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To attribute feature-flag API usage to cohorts in a Node.js service, record flag evaluations and configuration refreshes as separate events, stamp each one with a stable cohort key and the configuration version in effect, count failed and retried attempts, and split shared polling cost with a written rule. Keep cohort labels bounded in your metrics. OpenTelemetry’s default limit of 2,000 unique attribute combinations per metric stream means that an unbounded cohort label can cause overflow, and overflow drops the attributes you need to filter on.
Start with what your provider actually bills
“Feature-flag API request” is not a universal billing unit. PostHog’s “Cutting feature flag costs” documentation states the rule for server-side SDKs: “each call that evaluates flags, such as evaluateFlags(), evaluate_flags(), getFeatureFlag(), getAllFlags(), or isFeatureEnabled(), makes a request to the /flags endpoint and incurs a billable event unless local evaluation resolves it.” The same page documents polling of flag definitions as a separate charge when local evaluation is in use (PostHog, “Cutting feature flag costs”).
Two consequences follow for your design. First, the same flag check can be free, metered as an evaluation, or metered as part of a background poll, depending on SDK mode. Second, optional analytics events are not the billing basis: PostHog says its $feature_flag_called events are not what it bills on. Confirm the rules for your exact SDK version and plan before you map any internal event to an invoice line.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsSeparate evaluations from configuration refreshes
A single “flag cost” counter hides the question you most need answered: was the spend caused by application behavior or by keeping configuration current? Use three event families.
#1 Best Overall
| Event | Triggered by | Key outcome fields | Billing relationship |
|---|---|---|---|
flag_evaluation |
Application code asking for a flag value | Provider, SDK mode, flag key or bounded flag category, result, cohort, configuration version | Billable when the provider treats the call as an evaluation request; PostHog bills server-side calls unless local evaluation resolves them |
flag_config_refresh |
One poll or refresh attempt for flag definitions | Outcome (unchanged, changed, error, timeout, rate limited), HTTP status class, duration, configuration version or ETag | Billed separately where the provider charges for definition polling; PostHog documents this for local evaluation |
flag_config_refresh_retry |
A retry after a failed or limited refresh attempt | Attempt number, retry reason, backoff duration, retry-after bucket | Treat as a refresh attempt for capacity purposes; whether it is billed depends on the provider’s counting of failed requests |
You can model retries as their own event or as an attempt_number attribute on flag_config_refresh. The important rule is that every attempt is recorded, not only successful refreshes.
Tag every event with cohort and configuration context
A cost figure is only interpretable if you know which rules produced the behavior being measured. When a cohort sees a different flag result, the configuration version tells you whether it was served a different rule set or the same one. Include these fields on each event:
- provider and sdk_mode (for example, server-side evaluation or local evaluation), so billing classes stay separate.
- environment, so staging traffic does not mix into production cost.
- cohort_id, a stable pseudonymous key. Derive it with a keyed hash of the tenant or user identifier rather than sending the raw identifier to broadly exported metrics.
- config_version, using the provider’s ETag where one is returned, or a hash of the definition payload where it is not.
- outcome and http_status_class, so rate-limited responses are distinguishable from transport errors.
- observed_at and duration, for ordering and latency analysis.
- allocation_basis, the identifier of the shared-work rule applied to this record (see the next section).
Allocate shared polling cost with a written rule
One definition poll often serves many cohorts. Its cost has to be split somehow, and the split must be stated before you compare cohorts. Choose one rule and record its identifier in every derived record.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #2
| Rule | How it works | Strength | Weakness |
|---|---|---|---|
| Equal split | Each refresh’s cost is divided equally among cohorts it served | Simple to explain and reproduce | Penalizes small cohorts that evaluate rarely |
| Evaluation-volume split | Refresh cost follows each cohort’s share of observed evaluations in the same window | Tracks who actually consumes the configuration | Depends on complete evaluation counts; sensitive to traffic spikes |
| Direct assignment | A refresh is charged wholly to one cohort because a separate definition set is dedicated to it | Most accurate when configuration is truly cohort-specific | Only valid when the poll truly serves one cohort |
Keep raw attempt totals in a separate store. The allocated view is derived from them, so you can switch rules later without losing history. Any example figures in an internal review should be labeled as illustrations; they are not measurements.
Handle failed attempts, retries, and the last-known-good snapshot
Failed requests consume capacity and can create provider usage, so a dashboard that only counts successful refreshes understates both cost and load. Use this sequence in your refresh loop:
- Record every attempt with
attempt_numberandoutcomebefore deciding what to do next. - On a rate-limited or failed response, read the
Retry-Afterheader if the provider sends one and treat it as the minimum wait. Confirm the provider’s exact semantics in its API reference, because quota scope and retry behavior vary by product. - If no delay is supplied, use exponential backoff with bounded jitter and a fixed maximum attempt count. Do not raise retry frequency to recover faster; that increases the request budget you are trying to protect.
- Keep serving the last validated configuration. Validate each snapshot against a schema before swapping it in, and record
snapshot_agecontinuously. - Alert when snapshot age passes the maximum you set for your rollout risk, not only when requests fail.
Decide who owns the polling loop
When every Node.js process reacts independently to the same limit, the retries multiply. Three designs are common, and each moves the cost and failure boundary differently.
| Design axis | Central poller or shared cache | Per-process polling | Local evaluation with provider SDK |
|---|---|---|---|
| Request fan-out | Stays low when many processes share one source | Grows with process count | Depends on the provider’s polling and cache behavior |
| Configuration freshness | Set by one poll cadence and propagation path | Each process refreshes on its own schedule | Set by the SDK’s refresh policy |
| Failure boundary | The shared poller or cache becomes critical infrastructure | Failures are isolated per instance, but retries can multiply | The SDK handles some mechanics; you still need monitoring |
| Cohort attribution | Needs an explicit allocation rule | Direct when configuration is cohort-dedicated; otherwise shared | Evaluation and refresh charges must be separated by the provider’s definitions |
Vendor examples show how much these defaults differ. PostHog documents a 30-second default polling interval for feature-flag definitions, ETag requests for unchanged definitions, and sharing definitions across instances; its ETag support is tied to Node.js SDK 5.17.2 or later. It also warns against local evaluation in edge or Lambda-style environments where an instance may be initialized per invocation (PostHog, “Cutting feature flag costs”). Atlassian Forge’s server-side SDK keeps locally cached evaluations and polls for configuration updates every 60 seconds after initialization (Atlassian Developer, “Feature flags server-side SDK”; page last updated May 18, 2026). Neither default is universal.
Make the freshness-versus-budget tradeoff explicit
Every polling interval is a bargain between how quickly a change reaches a process and how many requests the fleet spends to learn about it. Work the arithmetic before you choose an interval.
- A 30-second interval produces 2,880 polls per day per process. Over a 30-day month, that is 86,400 polls per continuously running server. This matches PostHog’s stated arithmetic for unchanged polls, and PostHog adds 10 requests for each poll that returns new definitions. It is the vendor’s example calculation, not a measurement of your workload.
- Illustratively, 40 instances each polling every 30 seconds would issue 115,200 unchanged polls per day. Doubling the interval halves that figure, but a change can take up to twice as long to reach every instance.
- Atlassian’s 60-second update poll yields 1,440 polls per day per process under the Forge SDK’s stated behavior.
Use a decision order rather than a fixed default:
- Set the maximum acceptable propagation delay for a rollout change, based on the blast radius of a bad flag.
- Compute the poll count per day for your process count at that delay, then compare it against the provider’s quota and billing terms.
- If the budget cannot meet the delay, change the distribution boundary (a shared poller or a local-evaluation SDK) instead of increasing retry frequency.
- If the delay is too long for rollouts that must be fast, use a provider mechanism for targeted refresh, if one exists, rather than shortening the global interval.
Instrument Node.js telemetry before application code loads
OpenTelemetry’s Node SDK reference warns that initializing telemetry late can leave no-op implementations in place, so instrumented modules silently emit nothing. Start the SDK in its own file and preload it. OpenTelemetry’s JavaScript documentation lists traces and metrics as stable, and supports the active and maintenance LTS versions of Node.js (OpenTelemetry, “JavaScript”).
Rank #4
// instrumentation.js
const { NodeSDK } = require('@opentelemetry/sdk-node');
const sdk = new NodeSDK({ /* metric and trace exporters for your backend */ });
sdk.start();
// start the service with the SDK loaded first:
// node -r ./instrumentation.js server.js
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Define metrics with bounded cohort labels
Metrics aggregate by unique attribute combination, and each combination needs its own aggregation state. OpenTelemetry’s metrics documentation describes the default cardinality limit as 2,000 per metric stream, which can be overridden with a View. When the limit is reached, further measurements are folded into an overflow point that keeps the overall total but drops the original attributes, so a cohort-filtered query will undercount (OpenTelemetry, “Metrics”).
const { metrics } = require('@opentelemetry/api');
const meter = metrics.getMeter('feature-flags');
const evaluations = meter.createCounter('flag.evaluations');
const refreshes = meter.createCounter('flag.config.refresh');
const refreshLatency = meter.createHistogram('flag.config.refresh.duration', { unit: 's' });
const snapshotAge = meter.createHistogram('flag.snapshot.age', { unit: 's' });
// cohort_group is a fixed set such as 'control', 'treatment', 'internal', 'other'
evaluations.add(1, { provider: 'posthog', sdk_mode: 'local', cohort_group: 'treatment', result: 'on' });
refreshes.add(1, { provider: 'posthog', outcome: 'rate_limited', http_status_class: '4xx' });
Count the label combinations before you ship. For example, 3 providers × 2 SDK modes × 4 outcomes × 3 cohort groups × 5 status classes gives 360 combinations, well under the default limit. Putting a raw or pseudonymous cohort ID on the same metric can push a fleet past 2,000 combinations quickly. Keep high-cardinality cohort identity in logs or traces that support per-identifier lookup, and roll it up into cohort groups for metrics.
Validate the pipeline before using it for chargeback
Do not treat the dashboard as an invoice until these checks pass:
Best Value
- Exporter totals match application attempt counters for the same window.
- Provider usage reports match the request classes the provider actually bills, with evaluations and definition polls reconciled separately.
- Each evaluation record carries the cohort assignment and configuration version that were in effect when it was made.
- Refresh failures and retries line up with snapshot-age spikes and stale-evaluation counts.
- No overflow series appears in the cohort metrics; if one does, widen the View limit or reduce labels before reporting cohort totals.
These checks are a practical reconciliation pattern rather than a published cross-vendor standard. The provider documentation settles the billing distinctions, and the OpenTelemetry documentation settles the metric limits; the rest is your own control process.
Consider the reconciliation an ongoing control. Provider billing rules, default polling intervals, and SDK behavior change between releases, so recheck the provider’s current documentation and your installed SDK version whenever you change polling, upgrade the SDK, or change plans.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →

