Preprocessing
What you'll see here
Every transformation applied before training, in order. Nothing is imputed — the dataset has zero missing values.
Outlier audit — IQR method (Q1 − 1.5·IQR, Q3 + 1.5·IQR)
| Variable | Lower fence | Upper fence | Below | Above | % flagged | Treatment |
|---|---|---|---|---|---|---|
| no_of_employees | -2,701 | 7,227 | 0 | 1,556 | 6.11% | Retained; negatives clipped to 1, then log-transformed. |
| yr_of_estab | 1,933 | 2,049 | 3,260 | 0 | 12.79% | Retained; converted to company_age (capped at 0). |
| prevailing_wage | -76,565 | 218,316 | 0 | 427 | 1.68% | Retained; annualised by unit, then log-transformed. |
| yearly_wage (engineered) | -69,468 | 241,401 | 0 | 2,387 | 9.37% | Retained |
| company_age (engineered) | -32 | 84 | 0 | 3,260 | 12.79% | Retained |
Outliers are genuine business values (very large employers, high-wage roles, century-old firms), not data-entry errors — so they are retained rather than removed. Tree ensembles split on rank order and are insensitive to magnitude; log transforms further compress the tails. The only true error corrected is 33 rows with a negative employee count.
Pipeline steps
1. Drop identifier (case_id)
case_id is a unique identifier (25,480 unique values) with no predictive signal — dropped before any model training.
2. Clip negative employees
no_of_employees < 1 is clipped to 1 (employer must have staff).
3. Engineer company_age
company_age = max(0, 2016 − yr_of_estab). Captures firm maturity without leaking a raw year.
4. Annualize wage
yearly_wage = prevailing_wage × { Hour:2080, Week:52, Month:12, Year:1 }. Removes unit confusion.
5. Log-transform
log₁₊ applied to employees and yearly_wage — both right-skewed with a long tail.
6. One-hot encode categoricals
continent, education, region_of_employment, unit_of_wage → binary indicators.
7. Binary encode Y/N
has_job_experience, requires_job_training, full_time_position → 0/1.
8. Stratified split
70/30 train-test, seed=42, preserving denial rate in both partitions.
9. Resampling variants
Random oversampling of Denied and random undersampling of Certified — trained separately on Modeling → Resampling page.
AI: Preprocessing rationale
Click Generate to produce an evidence-based observation, insight, recommendation, and business benefit for this section.