← Back to Projects

Synthetic Data Analyzer

Generate synthetic tabular data, protect sensitive info, and measure how closely it preserves the original dataset.

PYTHONSTATISTICSPANDAS
Dashboard Configuration

The Problem

Organizations often need realistic datasets for development, testing, analytics, demonstration, and research. However, directly sharing real datasets exposes sensitive Personally Identifiable Information (PII) and risks violating privacy regulations.

This project focuses on generating synthetic data that is completely structurally and statistically similar to the original dataset, without containing any of the original sensitive information.

Data Analysis

DATASET ANALYSIS

Rows
10,000
Columns
18
Numeric
11
Categorical
7
PII Fields
4
Outliers
127
Missing Values
82

Synthetic Generation Pipeline

Processing Pipeline
ORIGINAL DATA
↓
STATISTICAL PARAMETERS
↓
GAUSSIAN COPULA
↓
SYNTHETIC DATA

No ML. No LLMs. Pure Classical Statistics.

The project uses mathematically rigorous, deterministic algorithms to preserve utility and enforce privacy.

Differential Privacy (Laplace Mechanism)

Adds controlled mathematical noise to statistical parameters (mean, std) ensuring strong privacy guarantees (epsilon = 1.0) against reconstruction attacks.

Auto-Outlier IQR Cleaning

Automatically mitigates heavy skew and capping outliers using Interquartile Range (IQR) bounds during analysis.

Smart PII Semantics (Faker)

Automatically detects columns like `first_name`, `email`, and generates hyper-realistic fake semantic replacements rather than randomized text.

Primary Key Safety & Sub-schemas

Intelligently sequences unique identifier columns (`ID-1`, `ID-2`) and groups related data logically for unmatched realism.

Quality Validation

Validation Results
Distribution Comparison

Preserving marginal distributions (shape) via KS-Test.

Correlation Comparison

Preserving feature correlations via Gaussian Copulas.

TECHNOLOGY

Python 3.13PandasNumpyScipyFakerFastAPIReact 18ViteTailwindRecharts