Home

Preprocessing

What you'll see here

Every transformation applied before training, in order. Nothing is imputed — the dataset has zero missing values.

Outlier audit — IQR method (Q1 − 1.5·IQR, Q3 + 1.5·IQR)
VariableLower fenceUpper fenceBelowAbove% flaggedTreatment
no_of_employees-2,7017,22701,5566.11%Retained; negatives clipped to 1, then log-transformed.
yr_of_estab1,9332,0493,260012.79%Retained; converted to company_age (capped at 0).
prevailing_wage-76,565218,31604271.68%Retained; annualised by unit, then log-transformed.
yearly_wage (engineered)-69,468241,40102,3879.37%Retained
company_age (engineered)-328403,26012.79%Retained

Outliers are genuine business values (very large employers, high-wage roles, century-old firms), not data-entry errors — so they are retained rather than removed. Tree ensembles split on rank order and are insensitive to magnitude; log transforms further compress the tails. The only true error corrected is 33 rows with a negative employee count.

Pipeline steps
1. Drop identifier (case_id)
case_id is a unique identifier (25,480 unique values) with no predictive signal — dropped before any model training.
2. Clip negative employees
no_of_employees < 1 is clipped to 1 (employer must have staff).
3. Engineer company_age
company_age = max(0, 2016 − yr_of_estab). Captures firm maturity without leaking a raw year.
4. Annualize wage
yearly_wage = prevailing_wage × { Hour:2080, Week:52, Month:12, Year:1 }. Removes unit confusion.
5. Log-transform
log₁₊ applied to employees and yearly_wage — both right-skewed with a long tail.
6. One-hot encode categoricals
continent, education, region_of_employment, unit_of_wage → binary indicators.
7. Binary encode Y/N
has_job_experience, requires_job_training, full_time_position → 0/1.
8. Stratified split
70/30 train-test, seed=42, preserving denial rate in both partitions.
9. Resampling variants
Random oversampling of Denied and random undersampling of Certified — trained separately on Modeling → Resampling page.
AI: Preprocessing rationale

Click Generate to produce an evidence-based observation, insight, recommendation, and business benefit for this section.