iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
You can combine Nomad with ephemeral CI runners and use Temporal as a possible orchestration layer, but the cited official documentation does not establish a ready-made Nomad–Temporal integration. The supported building blocks are clearer: Nomad can dispatch parameterized jobs, its separate Autoscaler can add client capacity, and GitHub recommends ephemeral self-hosted runners that handle one job and then deregister. Treat workflow state, retries, cancellation, and cleanup between those systems as design work to verify—not as guaranteed integration behavior.
What each part of the architecture does
Keep the responsibilities distinct. GitHub identifies work and routes it to an eligible runner. A Nomad parameterized job can start a separate dispatched job instance. Nomad places that work on available clients, while the Nomad Autoscaler can adjust task-group allocation counts or provision and decommission Nomad client nodes. A GitHub ephemeral runner is removed from GitHub after one job.
Temporal can be considered for coordinating the lifecycle, but the Nomad and GitHub documentation cited here does not define Temporal’s workflow or activity design, provide a native integration, or establish retry, cancellation, or cleanup guarantees for this combination. Before implementing that layer, verify its behavior against current Temporal documentation and test how it interacts with both Nomad and GitHub.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchSeparate the control plane from runner capacity
A useful design boundary is to keep CI event intake and orchestration separate from the worker capacity that executes untrusted or potentially changing job code. A request can lead to a Nomad dispatch for runner work; Nomad then schedules it where capacity exists. If there are no available clients, node scaling may be needed. Allocation scaling and client-node scaling are separate controls, so decide which one addresses each bottleneck.
#1 Best Overall
How a job can move through the system
- Accept a CI request. Define how the event reaches your control plane and how duplicate deliveries are recognized. The exact event bridge is an implementation choice; the cited documentation does not specify one for Temporal and Nomad.
- Dispatch runner work to Nomad. Nomad’s job dispatch command reference documents dispatching a parameterized job as a distinct job instance. The dispatch payload is limited to 16 KiB. With ACLs enabled, the caller needs the
dispatch-jobcapability in the namespace. An idempotency token is available; use it deliberately to help handle repeat dispatch requests. - Schedule and scale capacity. Nomad evaluates placement for the dispatched work. The separate Nomad Autoscaler can scale task-group allocations or Nomad client nodes. Nomad’s job scaling block provides a control interface for external autoscalers, whose policies consume its data.
- Run one job on an ephemeral runner. Configure the runner so it is registered as ephemeral and is not reused for another job. GitHub says an ephemeral runner handles one job and is automatically deregistered afterward. Forward its logs to external storage before relying on ephemeral capacity; deregistration otherwise makes later troubleshooting harder.
- Track completion and failure across systems. Decide how your orchestration layer learns that the CI job and Nomad allocation have ended, and what happens if either system stops reporting. Specify timeout, retry, cancellation, and cleanup behavior explicitly. Those behaviors are not established by the Nomad and GitHub references cited here.
Choose a scaling trigger that fits the workload
| Approach | What it means | Trade-off to assess |
|---|---|---|
| Webhook-driven scaling | A delivered CI event prompts your control plane to request or scale runner capacity. | GitHub notes that webhook delivery timeliness can affect autoscaling reliability. Account for delayed or repeated delivery and decide how requests are reconciled. |
| Scale-set-aware approach | Use an approach designed to respond to runner demand at scale. GitHub points to Actions Controller or the Scale Set Client for larger-volume scenarios. | Assess whether it fits your volume and event-handling needs. The cited documentation does not establish direct integration with Nomad or Temporal. |
Neither option removes the need to define how Nomad receives work, how capacity is made available, and how stale or failed requests are reconciled. Choose based on expected workload volume and acceptable queue delay, then test the complete path rather than assuming an event automatically produces a usable runner.
Plan for queues, access, and runner security
Restrict which workflows can reach self-hosted runners
Self-hosted runners execute workflow code in your environment. GitHub warns that untrusted code can compromise them, including code run by public-repository fork workflows. Use runner groups and access controls to limit which repositories and workflows can target the runner pool. Do not expose privileged infrastructure to untrusted jobs merely because the runners are short-lived.
Ephemeral runners reduce reuse between jobs, but that is not a substitute for access boundaries. Keep credentials and network access scoped to what a job requires, and make sure a runner cannot be selected by workflows that should not run on your infrastructure.
Make queueing and capacity exhaustion visible
GitHub’s runner reference says a job without a matching runner remains queued and fails if it stays queued for more than 24 hours. Set expectations for queue delay, alert on capacity exhaustion, and define a response for jobs that cannot acquire a matching runner. Labels and groups must align with the workflow’s runner requirements; otherwise extra Nomad capacity may not resolve a routing mismatch.
Rank #3
Design cleanup as an explicit path
Do not assume that cancellation in one system automatically stops work in the others. Test what happens when a GitHub job is canceled during runner startup, when a Nomad allocation fails, or when the orchestration service is unavailable. Establish who detects orphaned allocations and runner registrations, and how each is cleaned up. The available Nomad and GitHub references do not promise cross-system cancellation or cleanup.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Use the Nomad scaling tutorial as a pattern, not a production blueprint
HashiCorp’s on-demand batch cluster scaling tutorial demonstrates provisioning Nomad clients when batch work queues and decommissioning them after work finishes. HashiCorp explicitly warns that the demo has billable costs and is not suitable for production use as-is. Treat it as an example of a scaling pattern, not a complete production architecture for CI runners.
Rank #4
Before deploying, review the security model, cost controls, capacity limits, failure handling, and cleanup behavior for your environment. The tutorial itself calls for review of a production reference architecture; the fact that a demo provisions and removes clients does not establish that it is safe or complete for your workload.
Recommended Free Tools
Quick Recap
Implementation checklist
- Choose whether to scale task-group allocations, Nomad client nodes, or both, and identify the bottleneck each control is meant to address.
- Keep dispatch payloads within Nomad’s 16 KiB limit and grant dispatch callers only the required namespace capability.
- Define duplicate-event handling and use the dispatch idempotency token where appropriate.
- Use GitHub ephemeral runners for autoscaling, and forward logs externally before runner deregistration.
- Scope runner groups and repository access, especially where public-repository contributions may run untrusted code.
- Test delayed events, repeated events, no available Nomad clients, runner startup failure, job cancellation, and orchestration-service outages.
- Document queue-delay expectations, capacity alerts, retries, and orphan cleanup; verify Temporal-specific behavior against its current documentation.
- Review cost and production readiness rather than deploying HashiCorp’s tutorial demo unchanged.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

