Data Quality & EDA
Anomaly & Outlier Detection
Detect statistical anomalies and distribution distortions in your data streams.
Recipe Overview & Methodology
Identify abnormal events and measurement spikes before they corrupt downstream reports. Uses statistical IQR boundaries, rolling z-score deviations, and distribution kurtosis scans.
Statistical Validation Rules
- Tukey's Fences method with 1.5x and 3.0x Interquartile Range (IQR).
- Modified Z-Score using Median Absolute Deviation (MAD) for robust scaling.
- Temporal moving window anomaly filtering for time-series metrics.
- Automated alert trigger payloads for workflow pipelines.
Sample Output Metrics
Outlier Rate
1.2%
14 / 1,200 rows
Max Z-Score
+4.61
Extreme positive anomaly
Distribution Skew
1.84
Right-skewed tail
Recommended Action
Winsorize / Flag
Clean before modeling
Execution LogicClustey Engine
# Statistical Outlier Boundary Scan
q25, q75 = dataset['metric'].quantile([0.25, 0.75])
iqr = q75 - q25
lower_bound = q25 - (1.5 * iqr)
upper_bound = q75 + (1.5 * iqr)
anomalies = dataset[(dataset['metric'] < lower_bound) | (dataset['metric'] > upper_bound)]Ready to apply this recipe to your own dataset?
Connect your dataset and generate this analysis in one click.