Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Splink usually isn’t the thing that fails. It is a Python package, and it can be called from Python code inside a web service. What breaks is the fit between a full record-linkage run and the limits of a request/response path: how long the gateway waits, how much memory the function gets, and how the database connection is shared. The gateway often gives up before your code does, so the client sees a timeout while the job is still running. For short, predictable linkage runs, a synchronous endpoint can work. For larger or variable jobs, accept the request and process it in the background.
Identify which failure you have
“Can’t run in an API” covers several different problems, and each has a different fix. Start by matching your symptom to one of these categories.
| Symptom | Likely cause | First check |
|---|---|---|
| Import or deployment error, module not found | Dependencies not packaged for the runtime, or an unsupported Python version | Confirm Python 3.10 or later and the exact Splink version in the deployed environment. The getting-started guide lists 3.10+ as the requirement. |
| Client receives 504 or a timeout, but logs show the job still running | The gateway’s integration deadline is shorter than the job | Compare measured job duration with the gateway’s documented deadline for your API type. |
| Function timeout error in logs | The compute function’s own timeout was reached | Compare measured duration with the configured function timeout. On AWS Lambda, the ordinary maximum is 900 seconds. |
| Process terminated, out-of-memory message | Memory allocation too small for the input and the candidate pairs your blocking rules generate | Read peak memory from the logs for a realistic input, then raise the allocation or reduce the input per job. |
| Works for one request, fails under parallel requests | Shared database connection used across threads | Check whether all requests use one module-level DuckDB connection. |
| Database error with a non-default backend | Spark or PostgreSQL misconfigured or unreachable from the function | Test the backend connection from the same network context, outside Splink. |
| Slow first request, fast later requests | Cold start: package import and setup run on each new container | Time a cold invocation and a warm one separately. |
Know the limits you are working against
Three separate limits can end a request, and only one of them is set by your code.
- The compute timeout. AWS Lambda’s ordinary invocation timeout is configurable from 1 to 900 seconds (15 minutes). Specialised invocation modes and other AWS services can differ, so check the one you use.
- The gateway deadline. API Gateway applies its own integration timeout. The AWS documentation consulted for this article cites a 29-second value, but the applicable limit depends on API type, integration mode, and configuration, so confirm it for your deployment before relying on it.
- The client’s patience. A mobile app, browser, or upstream service may give up even earlier than the gateway.
Raising the function timeout cannot fix a shorter gateway deadline. The request has to finish inside the smallest of these three limits.
#1 Best Overall
Measure before you choose a design
Splink’s workflow has three stages: estimating model parameters, predicting matching pairs, and clustering the results. Its project repository says it can link a million records on a laptop in around a minute. That is the project’s own broad claim, not a guarantee for your data. Duration depends on row counts, column choices, blocking rules, the backend, and the compute you allocate, so measure your own workload.
- Reproduce the failure with a fixed input file and record the Splink version, Python runtime, backend, row and column counts, and memory and CPU allocation.
- Time each stage separately: parameter estimation, prediction, and clustering. Log the duration of each.
- Test with a realistic upper-bound input, not a small sample. Add headroom, because a sample can hide the cost of large candidate-pair sets.
- Compare the largest measured duration with both the function timeout and the gateway deadline. Use the smaller of the two as your budget, minus a safety margin.
- Check whether the input is reloaded on every request, whether the model is retrained when a saved model would do, and whether any intermediate results are rebuilt each time. These are hypotheses to test, not confirmed causes.
- Check memory at peak. An out-of-memory kill often appears as a generic failure with no timeout message.
Handle concurrency and shared connections
The DuckDB Python documentation warns that the duckdb module uses a shared global database, which can lead to hard-to-debug issues when used from multiple packages or threads. If your service runs Splink from a module-level connection, parallel requests can interfere with each other. Give each request, or each worker, its own deliberately managed connection object. Then test parallel requests explicitly, because a single-request test will not reveal this problem.
Choose between a synchronous endpoint and a background job
| Option | Fits when | Trade-offs |
|---|---|---|
| Synchronous endpoint | Input size is bounded, and measured upper-bound runtime sits well inside the gateway deadline | Simplest to build and call. The client waits, and a slower run than expected will fail the request. |
| Background job | Input size varies, or runtime is unpredictable or may exceed the deadline | More resilient and allows status reporting. Requires job state, a status endpoint, and a result path or notification. |
| Spark or PostgreSQL backend | Data already lives in that system, or a single process cannot handle the workload | Splink documents both as alternative backends, but neither is an automatic fix. Each adds infrastructure and its own failure modes. |
AWS publishes an asynchronous API processing pattern that follows this model. It is an AWS implementation, and the same architecture applies on other platforms with different services.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Build the background job pattern
- The client sends a POST request. The API validates the input, stores a job record with status queued, and returns HTTP 202 with a job ID.
- A message queue or job runner hands the job to a worker. The worker is not bound by the gateway deadline.
- The worker updates the job record as it moves through running, succeeded, or failed, and writes the result to storage the client can reach.
- The client polls a status endpoint using the job ID, or receives a notification when the job finishes.
- Set a retry and dead-letter policy so that failed jobs are visible and can be investigated rather than lost.
When a synchronous endpoint is still the right call
If your measured upper-bound run finishes comfortably inside the gateway deadline, and the input size is capped at the API layer, a synchronous endpoint is simpler and perfectly reasonable. Keep the concurrency fix from above in place, and add a check that rejects inputs above the tested size so that a larger request is routed to the background path instead of timing out.
Rank #3
Summary of the approach
Match the symptom to a category, measure each stage against the smallest applicable limit, fix shared connections, and then choose synchronous or background processing based on measured upper-bound runtime. Splink can run inside an API; the question is whether the request path gives it enough time, memory, and isolation.
Quick Recap
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

