Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To build and deploy a production-ready Node.js API on Cloud Run, make the server listen on Cloud Run’s injected PORT, deploy it from source or a controlled container image, and configure health checks, service identity, secrets, concurrency, and scaling for your workload. A successful deployment is only the start: verify the new revision’s health and traffic behavior before treating it as ready for production.
Prepare the project and deployment environment
Before deploying, select or create a Google Cloud project, install or update the Google Cloud CLI, authenticate, choose a region, and enable the APIs required by your deployment path. Choose a region with both user latency and the location of dependent Google Cloud resources in mind; also confirm that the services your API needs are available there.
IAM requirements depend on whether you deploy from source or an existing image, as well as your organization’s policies. The Google Cloud Node.js quickstart describes the required roles for its deployment path and says the build service account needs the Cloud Run Builder role. Check the current requirements for your chosen path rather than granting a broad set of permissions by default.
Make the Node.js server listen on Cloud Run’s port
Cloud Run provides the listening port in the PORT environment variable. Your application must bind to it rather than assuming a fixed port. The official Node.js quickstart uses this minimal Express pattern:
#1 Best Overall
const port = parseInt(process.env.PORT) || 8080;
app.listen(port, () => {
console.log(`Listening on port ${port}`);
});
The fallback to 8080 is useful for local development; Cloud Run supplies the port when it runs the container. This small example demonstrates port binding only. It does not establish that an API is secure, healthy under load, or production-ready.
Choose a deployment workflow
| Workflow | Build and image control | Good fit when |
|---|---|---|
| Deploy from source | Cloud Run builds a container image from the project source; source deployment automates the build step and uses a Dockerfile automatically. | You want a direct path from a project directory to a deployed service and do not need to manage the image build separately. |
| Deploy a container image | Your team builds and pushes an image to Artifact Registry, then deploys that image. This gives the team explicit control over the image supplied to Cloud Run. | Your release process requires an explicit image build, image review, or promotion between environments. |
Deploy from source
From the project directory, run the source-deployment command:
gcloud run deploy --source .
The command may prompt for a service name, region, API enablement, and whether to allow public access. Allow unauthenticated access only if the API is intended to be public. For a private service, configure authentication and test it using Cloud Run’s private-service flow.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #2
Deploy an image
For an image-based release, build and push the container image to Artifact Registry, then deploy that image to Cloud Run. This is a different workflow from --source: the team owns the image-build and promotion steps rather than asking Cloud Run to build from the source directory as part of deployment.
Set up health checks and verify revisions
Configure a startup health check so Cloud Run can determine when the container is ready to receive traffic. For an HTTP probe, the application needs to serve the configured probe path over HTTP/1. A successful startup probe indicates that the container is ready. If the default deployment startup health check fails, Cloud Run marks the new revision unhealthy and does not route traffic to it.
Cloud Run also documents startup, liveness, and readiness probe behavior. The configuration reference labels readiness probes Preview; check the current feature status and any applicability limits before relying on them. Do not assume that readiness probes are generally available.
Rank #3
Each configuration change creates a new, immutable revision. After deployment, check the new revision’s health and confirm which revision is receiving traffic. A deployment command completing is not, by itself, confirmation that the intended revision is healthy and serving requests.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteUse a dedicated service identity and manage secrets
Give the Cloud Run service a dedicated service account with only the permissions the API needs to call Google Cloud services. Google’s configuration guidance recommends minimizing permissions. Keep API keys, passwords, certificates, and other sensitive values in Secret Manager; do not commit them to source control or use build-time environment values as a shortcut.
Grant the service identity the Secret Manager Secret Accessor role for the required secret. Choose how Cloud Run delivers the secret based on when updates must become visible:
Rank #4
| Delivery method | What happens when the secret changes | Operational implication |
|---|---|---|
| Secret volume | The current secret value is fetched when the application reads it. | Can work with secret rotation; application behavior depends on when and how it reads the mounted value. |
| Environment variable | The secret value is resolved when an instance starts. | Running instances do not receive a changed value merely because the secret changed. Google recommends pinning environment-variable secrets to a specific version rather than latest. |
Harden the container where the application allows
Run the container as a non-root user when the application’s file access and runtime requirements permit it. Check dependencies for compatibility with Cloud Run’s execution environment: for example, Cloud Run execution can fail for setuid binaries, so do not assume that every general-purpose container behaves unchanged.
Set concurrency based on the API’s behavior
Concurrency controls how many requests Cloud Run can send to one instance at the same time. The current Google Cloud concurrency documentation, accessed October 7, 2026, sets the platform maximum at 1,000 concurrent requests per instance. Its documented defaults differ by deployment method:
Recommended Free Tools
| Deployment method | Documented default for a newly created service |
|---|---|
| Cloud Run console | 80 concurrent requests |
| Google Cloud CLI or Terraform | 80 times the number of vCPUs |
These are platform defaults, not recommended values for every API. Google’s documentation says, “Node.js is inherently single-threaded.” Asynchronous I/O can still let a Node.js process handle concurrent work, but CPU-bound handlers and shared mutable state need careful consideration. Validate that handlers and dependencies behave safely when requests overlap before raising concurrency.
- Higher concurrency can let fewer instances serve the same request volume and may reduce cost, provided the application handles parallel requests efficiently.
- Lower concurrency can provide more isolated scaling for workloads that need it, but can require more instances for the same request volume.
- A concurrency setting of one can impair scaling performance during traffic spikes.
Load-test representative traffic and observe CPU, memory, latency, errors, and instance counts before changing the setting. Treat the result as a workload-specific choice, not a general Node.js tuning rule.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Balance warm capacity, latency, and cost
Cloud Run scales instances in response to incoming requests. With zero minimum instances, an API can scale down to zero; a later request may experience the delay associated with starting an instance. Setting a minimum keeps a configured floor of instances warm and can reduce that scale-from-zero delay, but adds billing cost.
Minimum instances are not guaranteed capacity. Google’s documentation describes them as a best-effort target: capacity issues, rebalancing, crashes, quota limits, or billing issues can leave fewer healthy instances than configured. Google suggests considering at least three minimum instances for high availability, but that is guidance, not an availability guarantee.
Decide whether warm capacity is worth its baseline spend based on the API’s latency needs and workload. Estimate cost using current Cloud Run pricing and your own assumptions about request volume, resource allocation, and time warm; there is no meaningful cost figure without those inputs. Request timeout, CPU, memory, maximum instances, and minimum instances are separate controls, so tune each for the application rather than treating one setting as a substitute for the others.
Quick Recap
Verify the production deployment
- Confirm that the service listens on the injected
PORT. - Check that startup health succeeds and the intended revision receives traffic.
- Verify that public or authenticated access matches the API’s intended audience.
- Confirm that the service identity can access only the secrets and Google Cloud resources it needs.
- Observe request latency, errors, CPU, memory, and instance counts under representative traffic before settling on concurrency and scaling settings.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

