The hard part was not training a classifier. It was building a system where data identity, policy search, held-out evidence, serving behavior, and the authority to reject a result could be inspected together.

The repository began with a candidate pre-screening framing. Its data, however, is UCI Adult, a 1994 Census-income benchmark rather than applicant or job-performance records. I kept the difficult engineering question and rebuilt the surrounding system: what evidence should exist before a model score is allowed to become policy?

I rebuilt the project around that idea. It is now an auditable policy lab for tabular classification on the UCI Adult benchmark, with Logistic Regression, Random Forest, and XGBoost as interchangeable scoring models. The model is only one component. Every run must also explain which records it used, how it chose a decision policy, what the held-out data says, whether the governance gate accepts the evidence, and what an operating system would monitor.

SYSTEM MAP / TRUSTWORTHY MLThree planes keep a fairness audit inspectable.The workflow binds data identity, policy search, and serving evidence without turning an offline intervention into an unreviewed service rule.
  1. Data contract

    What the run may read and how every partition remains identifiable.

    BOUND INPUT
    1. Raw auditCanonical records retain the protected attributes needed for measurement.
    2. Digest-bound quality sidecarQuality findings and the model-ready table resolve to the same data identity.
    3. Exact 11-feature boundaryOnly declared prediction fields enter the model; audit fields stay outside.
    4. Official splitThe source test partition stays sealed while fit and validation remain distinct.
  2. Policy lab

    Where scores become candidate decisions, with sensitivity checks before selection.

    SEARCH + STRESS
    1. ScoresThe frozen model emits probabilities before any threshold is applied.
    2. Exhaustive two-threshold Pareto frontierEvery allowed threshold pair is evaluated against utility and disparity objectives.
    3. Global review bandOne shared uncertainty band routes borderline decisions for review.
    4. Intersectional and overlap sensitivityIntersections, validation retuning, and fixed-policy held-out overlap receive separate checks.
  3. Governance and operations

    What may pass, load, serve, and remain observable after the audit.

    VERIFY + OBSERVE
    1. Uncertainty-aware gateConfigured constraints and intervals own the verdict, not one point estimate.
    2. Seven-file integrity-bound bundleThe manifest binds every required artifact and refuses a mismatched load.
    3. Strict API simulationSchema, domains, categories, and policy behavior are exercised as service requests.
    4. Aggregate monitoringOperational rates surface drift and review pressure without row-level exposure.
END-TO-END CONTROL SPINEOne trace crosses all three planes.
  1. Dataset identity
  2. Quality evidence
  3. Selected policy
  4. Gate verdict
  5. Operational signals

Read the published result first

The reference study repeats the complete XGBoost workflow across five seeds. It does not select the friendliest run. Each run rebuilds the fit and validation split, searches the policy, evaluates the fixed result on the same official test partition, and then enters a strict comparability check before aggregation.

COMPARABLE RUNS5Seeds 0 through 4
POLICY CANDIDATES10,201Threshold pairs per seed
PAIRED BOOTSTRAPS500Per seed
GOVERNANCE0 / 5Bundles accepted
MODEL-READY ROWS45,222From 48,842 raw rows
HELD-OUT OVERLAP3,581Exact feature identities isolated

Across the five seeds, median adjusted accuracy was 0.8689 with a range of 0.8649 to 0.8709. Median disparate impact was 0.3657 with a range of 0.3142 to 0.4146; median selection-rate difference was 0.1754 with a range of 0.1236 to 0.1808; and the median signed true-positive-rate gap was 0.0368 with a range of -0.0233 to 0.0849. Every report was structurally valid, and every report was rejected by the configured gate. That separation between a valid artifact and an unacceptable policy result is the central outcome.

The five seeds reuse the same official test partition. Their spread measures sensitivity to the fit-validation split, model fitting, and policy selection, not five independent samples from a target population.

The benchmark also contains substantial identity repetition. 3,581 of 15,060 held-out rows share the exact 11-feature identity of a fitting or validation row. The fixed-policy novel-only view keeps 11,479 rows and reports it beside the primary evaluation rather than silently changing the official partition.

Start with an exact data boundary

The pipeline exposes an exact 11-feature input contract. fnlwgt, the target, the source split marker, and the protected fields used for auditing do not enter the model. Sex and binary race remain available to the evaluation path so disparity can be measured without leaking those fields into prediction.

Preparation begins with a raw-data audit. Missing-value handling, duplicate identities, repeated feature vectors, conflicting labels, and cross-split overlap are recorded in a quality sidecar. A cryptographic digest binds that sidecar to the model-ready table, so a report cannot quietly describe a different dataset than the one used by the run.

The source-defined test partition remains intact. The source training partition is divided into fitting and validation sets, and each has one job:

Partition Responsibility
Training Fit preprocessing and the selected scoring model
Validation Search thresholds, select the review band, and test policy sensitivity
Official test Measure the frozen model and frozen policy once

This separation matters because the threshold pair is a learned policy. Treating it as a harmless post-processing constant would create the same leakage problem as tuning model parameters on the test set.

Search a policy, not a convenient threshold

After the model is frozen, the validation scores enter an exhaustive two-dimensional search over protected-group thresholds. Every permitted threshold pair is evaluated, dominated candidates are removed, and the remaining Pareto frontier exposes the tradeoff between utility and disparity instead of hiding it inside one hand-tuned objective.

The selected offline policy also carries one global review band around uncertainty. The band depends only on model probability, not group membership. It routes borderline cases away from an automatic outcome and makes review load a measurable part of the policy.

Policy search remains validation-only. The official test set receives the already selected thresholds and review band without another optimization pass.

Ask two different overlap questions

Exact feature identities can recur across partitions in tabular benchmarks. The system treats that as evidence to investigate, not a reason to silently rewrite the official split.

First, an overlap-excluded validation sensitivity removes validation identities already present in training and reruns policy selection. This asks whether the chosen policy depends on repeated identities during tuning.

Second, the held-out sensitivity keeps the fitted model and policy fixed, then separates test records into overlapping and novel identities. This asks whether the reported evaluation changes when only unseen feature identities are considered. Because neither the predictions nor the policy are refit, the check stays diagnostic rather than becoming a second test-set optimization.

Make uncertainty part of the evidence

The report goes beyond aggregate sex metrics. It includes sex-by-race intersections, weighted and unweighted views, calibration diagnostics, paired bootstrap comparisons, and uncertainty intervals for disparity and utility measures.

Each subgroup receives an evidence state based on effective support. Sparse cells are marked as insufficient rather than converted into confident-looking point estimates. Worst observed spans remain visible, which prevents a favorable aggregate from erasing a smaller intersection with weaker behavior.

Repeated-seed stability is part of the implementation. The study runner executes the same resolved configuration across seeds, verifies that gate thresholds and policy settings remain comparable, and summarizes the distribution of decisions and diagnostics. Incomparable runs are rejected instead of being averaged together.

Keep the audit policy and service boundary distinct

The offline policy may use protected-group thresholds because its purpose is to study a fairness intervention. The service simulation has a different contract: protected attributes are absent from its request schema and never reconstructed inside the API.

Boundary Inputs Decision behavior
Offline audit Predictors plus protected fields held outside the model Applies the frozen group-specific thresholds and the shared review band for evaluation
FastAPI simulation Exactly the 11 declared predictor fields Applies the persisted protected-attribute-free service policy and returns auto_negative, manual_review_required, or auto_positive

The API rejects extra or missing fields, out-of-domain numbers, and unknown categories. It validates the fitted preprocessing vocabulary against the manifest before serving and refuses an integrity mismatch. The normal load path also refuses a bundle whose governance result is not accepted.

This boundary is deliberate. An offline fairness intervention cannot become live service behavior simply because both artifacts happen to sit in the same directory.

Let governance reject the result

The acceptance policy is persisted with the report. It evaluates utility, disparity, uncertainty, evidence sufficiency, and whether validation found a feasible policy candidate. Those thresholds are part of the run record, so rerunning the gate uses the same contract instead of whatever defaults happen to exist later.

All five reference runs were rejected by the configured governance gate. That is informative system behavior: progress on one measure cannot override a violation elsewhere, and an infeasible frontier cannot fall back to a baseline policy that appears to pass. The verdict travels with each bundle and controls whether the standard service path may load it.

Bind the implementation to its evidence

A completed audit writes a seven-file bundle through atomic publication:

Artifact Role in the trace
manifest.json Binds file digests, versions, feature contract, data identity, and governance state
model.joblib Stores fitted preprocessing and the scoring model
policy.json Versions the offline evaluation policy and protected-attribute-free service rule
predictions.csv Preserves paired scores and decisions for inspection
report.json Records protocol, metrics, uncertainty, sensitivity checks, and gate result
audit.html Provides a self-contained, print-ready engineering report
monitoring.json Captures the aggregate operational baseline

Loading is fail-closed. Required files, digests, schemas, feature order, categorical vocabulary, policy fields, and report bindings must agree before the bundle is considered usable.

The monitoring path works from aggregate counts rather than row-level records. It compares outcome mix, review pressure, score distribution, subgroup rates when available, and data-quality signals, then returns PASS, FAIL, or INSUFFICIENT_EVIDENCE. Derived rates are checked against their counts so a malformed snapshot cannot change the verdict unnoticed.

What the rebuild demonstrates

The project now connects ML experimentation with the engineering controls that make a consequential decision workflow reviewable:

  1. Data provenance is part of the executable contract.
  2. Policy optimization is isolated from held-out measurement.
  3. Fairness diagnostics include uncertainty, intersections, and overlap sensitivity.
  4. Offline protected-group analysis does not leak across the service boundary.
  5. Governance can stop artifact loading, and that refusal is preserved as evidence.
  6. Repeated-seed stability and aggregate monitoring share the same policy trace.

Charles Santhakumar and I built the original shared system at the University of Helsinki. I designed and implemented this expanded engineering architecture, audit workflow, governance layer, service boundary, and portfolio case study.

Inspect the evidence-bound source →

Open the five-seed evidence →