What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
First identify whether the failure is an orphaned foreign key or a relationship that is technically valid but unrealistic. Referential integrity means each non-null child key matches a parent key; relational quality also depends on details such as child counts, bridge-table rules, and which kinds of parents can have which children. Fix the schema or generation process that causes the specific failure, then validate each generated batch before loading or sharing it.
How do I preserve relationships between tables in synthetic data?
Give the generator an accurate description of the tables and their connections, and use a generation strategy that can account for those connections. Multi-table metadata describes tables and key relationships; the SDMetrics Multi Table Metadata guide uses parent-and-child tables as examples. Check that the metadata matches the actual schema rather than assuming that a generator can infer every relationship from column names.
Three different problems can look like “relationships are broken”:
- Referential-integrity failure: a non-null foreign-key value in a child table does not occur in the related parent table’s primary-key column.
- Relationship-description failure: the generator’s metadata has the wrong tables, key columns, data types, or key mapping—or does not describe the relationship at all.
- Relational-quality failure: keys resolve, but the generated pattern is implausible, such as an unrealistic number of children per parent or invalid parent-child combinations.
These require separate checks. A successful join only shows that keys can match; it does not prove that the generated relationships resemble the patterns needed by your application.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
Why are my synthetic foreign keys orphaned?
A common cause is generating related tables independently. If a child-table generator does not know which parent keys were generated, it can emit keys that do not exist in the parent table. A simple multi-table baseline can fail too: SDGym’s documentation says its MultiTableUniformSynthesizer randomly generates ID columns and does not ensure valid connections or referential integrity.
Before changing the model, check the source data and the metadata it consumes. Profile the real tables for duplicate parent keys, orphan child keys, nulls, inconsistent key types, and unexpected duplicates in bridge tables. Confirm that the declared primary and foreign keys match the database schema. SDV’s database-connector documentation describes schemas—including column names, types, and table connections—as information used to create metadata. Its AI Connectors bundle is an Enterprise feature, so that ingestion path is not available in every installation.
Use SQL checks against the generated tables to locate the failure. For example, with orders.customer_id referencing customers.id:
Rank #2
-- Non-null child keys with no matching parent key
SELECT o.customer_id, COUNT(*) AS affected_orders
FROM orders AS o
LEFT JOIN customers AS c ON c.id = o.customer_id
WHERE o.customer_id IS NOT NULL
AND c.id IS NULL
GROUP BY o.customer_id;
-- Duplicate parent keys
SELECT id, COUNT(*) AS occurrences
FROM customers
GROUP BY id
HAVING COUNT(*) > 1;
-- Null child keys, checked separately from orphan references
SELECT COUNT(*) AS null_customer_ids
FROM orders
WHERE customer_id IS NULL;
Adapt table and column names to your schema. The null check matters because the SDMetrics ReferentialIntegrity metric counts missing foreign-key values as valid; that behavior may not match a mandatory-key policy in your application.
Free tools Windows power users keep installed
One-click scans. No signup required.
How should I handle business constraints in the source data?
Decide whether each rule is truly universal before encoding it as a constraint. SDV’s troubleshooting guidance says, “A constraint should describe a rule that is true for every row in your real data.” If the input contains a row that violates a declared constraint, SDV reports a ConstraintsNotMetError.
You can remove a constraint or clean rows that violate it, but cleaning can make synthetic data less representative of the original input. If an exception is legitimate, model the exception rather than silently discarding it. Apply deterministic constraints only after checking the source records they are meant to describe.
Which generation strategy should I use?
For connected tables, prefer a relational or multi-table synthesizer that consumes relationship metadata and supports the rules your schema needs. SDV documents multi-table relational generation alongside evaluation and constraints. That is a documented capability, not proof that it will outperform every alternative or preserve every useful pattern in your specific data.
Where you need more specialized rules, SDV’s Constraint Augmented Generation documentation describes multi-table constraints including ForeignToPrimaryKeySubset, CompositeKey, and UniqueBridgeTable. CAG is described as an SDV Enterprise bundle; confirm that the feature is available in your edition and installed version before designing around it.
If separate generation is an engineering requirement, make the dependency explicit: establish the parent key set first, then assign child foreign keys from that set. Sample child counts and conditional values according to the relationship behavior your use case requires. This is a custom fallback, not a guarantee of realistic higher-order relationships; validate its output just as you would a synthesizer’s.
| Approach | What it can address | What to verify |
|---|---|---|
| Relational or multi-table synthesizer | Uses relationship metadata to generate connected tables; some products or editions offer additional multi-table constraints. | Whether the installed product supports your keys and constraints, and whether output preserves the distributions your use case needs. |
| Custom staged pipeline | Can establish parent keys before assigning child references and encode explicit sampling rules. | Key validity, child-count and conditional distributions, bridge-table rules, and any behavior not captured by the staged logic. |
Choose based on schema complexity, composite keys and bridge tables, required cardinality rules, scale, runtime, integration limits, and feature availability. No single approach is established as a universal fix.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How do I validate referential integrity in synthetic data?
Run checks on every generated batch after sampling and before loading or sharing it. Keep the failure counts and representative examples in a small report so a passing database load does not stand in for data-quality validation.
- Primary-key uniqueness: verify that generated parent keys are unique.
- Foreign-key resolution: check that each non-null child key occurs in the corresponding generated parent key set. SDMetrics ReferentialIntegrity measures the proportion of synthetic foreign-key values found in the synthetic primary-key column; missing values count as valid in this metric.
- Null policy: separately check null foreign keys wherever the relationship is mandatory.
- Cardinality: inspect child-row counts per parent against plausible bounds and business expectations. SDMetrics Diagnostic includes CardinalityBoundaryAdherence as a connection diagnostic.
- Bridge and business rules: check composite-key uniqueness, allowed parent-child combinations, and domain constraints that matter to downstream joins or tests. SDV’s CAG documentation lists constraint types such as CompositeKey and UniqueBridgeTable.
- Structure: compare generated table and column structure with the expected schema; SDMetrics Diagnostic includes table-structure measurements.
Do not fix an orphan by replacing its foreign key with an arbitrary valid ID. That can make the join succeed while attaching the child to the wrong parent. If repair is necessary, use a deterministic, auditable mapping and re-check the affected child distributions.
Best Value
How do I tell whether the relationships are realistic?
After key and schema checks pass, compare relationship patterns that matter to the consumers of the data: children per parent, combinations of parent categories and child types, and duplicate behavior in bridge tables. Set acceptance criteria for those patterns based on the downstream use. A database load succeeding establishes compatibility with that load path, not statistical utility.
Assess privacy separately from relationship quality. SDMetrics detection metrics ask whether a classifier can distinguish real from synthetic data, but its single-table detection guidance says not to use these metrics on primary- or foreign-key ID columns. It also cautions that a perfect score may indicate copied data and possible privacy leakage. A detection score therefore cannot replace relational-quality checks or a broader privacy assessment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

