Generate synthetic tabular data, protect sensitive info, and measure how closely it preserves the original dataset.

Organizations often need realistic datasets for development, testing, analytics, demonstration, and research. However, directly sharing real datasets exposes sensitive Personally Identifiable Information (PII) and risks violating privacy regulations.
This project focuses on generating synthetic data that is completely structurally and statistically similar to the original dataset, without containing any of the original sensitive information.

The project uses mathematically rigorous, deterministic algorithms to preserve utility and enforce privacy.
Adds controlled mathematical noise to statistical parameters (mean, std) ensuring strong privacy guarantees (epsilon = 1.0) against reconstruction attacks.
Automatically mitigates heavy skew and capping outliers using Interquartile Range (IQR) bounds during analysis.
Automatically detects columns like `first_name`, `email`, and generates hyper-realistic fake semantic replacements rather than randomized text.
Intelligently sequences unique identifier columns (`ID-1`, `ID-2`) and groups related data logically for unmatched realism.


Preserving marginal distributions (shape) via KS-Test.

Preserving feature correlations via Gaussian Copulas.