FINANCIAL SERVICES / CONSUMER LENDING

AnalyticsData Engineering

Automobile-Loan First-EMI Default Strategy Portfolio

A credit-policy analyst can test where to draw a first-EMI risk line, see the confidence interval around the answer and who it declines, and take a recommendation or a refusal to governance. On the published assumptions the honest output is a refusal: no evaluated band clears zero.

Automobile-Loan First-EMI Default Strategy Portfolio project cover

STAKEHOLDER VIEW

What this project is for.

Problem
A credit-policy analyst can test where to draw a first-EMI risk line across 233,154 real vehicle loans, see the confidence interval around the answer and which states it declines, and drill to the individual loans behind any number.
Intended user
A credit-policy analyst preparing a band recommendation for governance review.
Decision supported
Recommend or refuse one retrospective first-EMI policy band for governance review.
Outcome
Across 233,154 Indian vehicle loans, none of the three evaluated first-EMI risk bands clears zero at the published assumptions. The best, a conservative 16% cut-off, admits 44.69% of applications at a 15.29% observed default rate and lands at −₹0.07 cr once its 27,585 manual reviews are costed; resampling puts it above zero in only 39% of draws. That conclusion emerged from correcting a data defect: the bureau column encoded 'no score' as a number, and fixing it moved the same band from +₹1.52 cr to negative. But the refusal is conditional. A 21.7% first-EMI rate cannot be pure credit risk, and loans that missed because a mandate failed carry no credit loss, so the workbench reports how much of the label would have to be operational for each band to break even. Conservative needs 0.32%. The output is therefore a threshold a lender can test against its own payment plumbing, not a verdict on the loans.
What to try
Open the workbench and filter the October holdout cohort
Important limitation
The outcome is first-EMI default, not fraud and not long-horizon credit performance; the source has no fraud flag and no channel field.

Recommending a conservative first-EMI band

  1. 01Open the workbench and filter the October holdout cohort
  2. 02Compare the conservative band against the reference band on admitted volume, observed default rate and manual-review load
  3. 03Drag loss severity from 65% to 75% and watch every band go negative, and the recommendation become a refusal
  4. 04Read where the selected band breaks even on each assumption, and check the 27-cell grid behind the headline
  5. 05Open the KPI dashboard, drill into the loans behind the highest-volume state code, and export the monthly KPI pack
Technical reviewArchitecture, evidence, controls, deployment, and trade-offs

ARCHITECTURE

One source, one rendered system view.

One SQLite star schema feeds every surface through a shared parameterized SQL layer, so the API, the React application, the KPI pack and the published evidence cannot drift apart. A calibrated challenger is selected on calibration error rather than ranking alone, because a policy band needs probabilities that mean what they say — and the monitoring view then shows that on this holdout they do not, since observed default exceeds predicted in every decile. The economics are exposed as sliders rather than asserted, because the band ranking flips inside the plausible range of its own assumptions. Headline figures carry bootstrap intervals because the net result is a small difference between two much larger numbers. The whole platform runs as one process in one scale-to-zero container, with hand-authored inline SVG charts so no charting library ships to the browser.

Automobile-Loan First-EMI Default Strategy Portfolio system architecture
Pan to inspect the system; zoom controls are available when detail is needed.

End-to-end flow

  1. 01Ingest and validate the licensed Kaggle vehicle-loan archive
  2. 02Build a SQLite star schema and score every vintage
  3. 03Train a logistic baseline and a calibrated challenger
  4. 04Evaluate policy bands on a retrospective October holdout, with bootstrap intervals
  5. 05Cost the manual reviews each band generates, and stress-test across a 27-cell assumption grid
  6. 06Report who each band declines, and monitor population stability against predicted-versus-observed default

Technology stack

Microsoft Azure
Azure Container Apps
PythonJavaScriptscikit-learnSQLiteFastAPIReactDocker

Technology decisions

DecisionWhyAlternativeTrade-off
SQLiteA 233,154-row star schema fits comfortably in a single file, ships inside the container, and needs no separate database service on a free tier.PostgreSQL on a managed free tierNo concurrent writers and no network access to the store; acceptable because this workload is read-only and single-container.
A shared parameterized SQL layerThe API, dashboard, and KPI pack all read the same .sql files, so the evaluated evidence and the reported figures cannot diverge.Per-surface query codeLess freedom to shape queries per surface, in exchange for figures that reconcile by construction.
Dash mounted in-process via a2wsgiOne process fits one free container, and the dashboard becomes a same-origin path rather than a second host.Running the dashboard as a second serviceThe two apps share a process and a failure domain.

Evaluation and evidence

MetricPlain-language meaningScore / valueDataset / scenarioThreshold / baselineInterpretationEvidence
Net contribution, conservative bandThe only evaluated band that pays for itself once its own operating cost is counted.−₹0.07 crAugust–October 2018 retrospective first-EMI defaultNo threshold or baseline recorded.The outcome is first-EMI default, not fraud and not long-horizon credit performance; the source has no fraud flag and no channel field.economics.sensitivity
Assumption combinations the conservative band winsThe recommendation is robust to which band, and fragile as to whether to lend at all.22 of 27August–October 2018 retrospective first-EMI defaultNo threshold or baseline recorded.The outcome is first-EMI default, not fraud and not long-horizon credit performance; the source has no fraud flag and no channel field.economics.sensitivity
Decile calibration errorWhy the calibrated challenger was selected over a better-known baseline: a policy band is only as trustworthy as its probabilities.0.04542August–October 2018 retrospective first-EMI defaultNo threshold or baseline recorded.The outcome is first-EMI default, not fraud and not long-horizon credit performance; the source has no fraud flag and no channel field.evaluation.holdout
Evaluation reproducibilityEvery published figure can be regenerated rather than taken on trust.byte-identicalartifacts/strategy_summary.json hashes to 03ee84454ef491b949b4c8d7ec41cb1eab063bbd049b7734b09fb7c5e08cdfb6 across independent runs.No threshold or baseline recorded.The outcome is first-EMI default, not fraud and not long-horizon credit performance; the source has no fraud flag and no channel field.evaluation.reproducibility

These are Retrospective holdout on August–October 2018 disbursements, with a 27-cell assumption sensitivity grid results, not a production service-level objective. Unknown values are shown as “Not recorded”; units and special characters retain their source meaning.

Technical terms and value conventions

TermPlain-language useHow this project uses it
First-EMI defaultThe borrower missed the very first monthly instalment.The only outcome this dataset records, and the only thing any band here is measured against.
AUROCHow well the model ranks riskier borrowers above safer ones; 0.5 is a coin flip.Reported at 0.64363, which is modest and stated as such rather than dressed up.
Brier score and ECEWhether a predicted 20% risk really happens about 20% of the time.Why the calibrated challenger was selected over the baseline; a policy band is only as trustworthy as its probabilities.
Manual-review rateThe share of applications a band sends to a human instead of deciding automatically.Treated as a capacity constraint, because a band that reviews everything is not operable.
Assumption-led contributionA modelled money figure built on stated assumptions, not measured profit.Shown with its 12% contribution rate, 65% loss severity and ₹1,500 review cost printed beside it, and adjustable, so it is never mistaken for observed profit and loss.
Break-even assumption valueHow far one assumption can move before the answer flips.The conservative band survives until loss severity reaches 0.693 against the 0.65 assumed — about 6.6% of headroom.

Data boundary

ClassificationSourcePermitted useExcluded data
publicVehicle Loan Default Prediction (Kaggle, avikpaul4u), local archive SHA-256 42573eba1d33248273cc2b098d96375cc6540fd4e1e4c01b2fabe7054779571bRetrospective analysis, demonstration, and public showcasingUNIQUEID; DATE_OF_BIRTH; CURRENT_PINCODE_ID; EMPLOYEE_CODE_ID; DISBURSAL_DATE; AADHAR_FLAG; PAN_FLAG; VOTERID_FLAG; DRIVING_FLAG; PASSPORT_FLAG; MOBILENO_AVL_FLAG

This dataset measures first-EMI default, not fraud. The source is a public Kaggle archive of Indian vehicle loans whose published licence is recorded as unknown and which is used here on the repository owner's authority; no confidential, client, or personal data is involved, and direct identifiers are excluded from every response. Nothing here makes an approval, denial, pricing, underwriting, or automated decision, and no fraud-detection capability is claimed.

Security and privacy controls

ControlImplementationEvidenceLimitation
Public data boundaryUNIQUEID; DATE_OF_BIRTH; CURRENT_PINCODE_ID; EMPLOYEE_CODE_ID; DISBURSAL_DATE; AADHAR_FLAG; PAN_FLAG; VOTERID_FLAG; DRIVING_FLAG; PASSPORT_FLAG; MOBILENO_AVL_FLAGsecurity.dependencies, security.injection, disclosure.scopeThis dataset measures first-EMI default, not fraud. The source is a public Kaggle archive of Indian vehicle loans whose published licence is recorded as unknown and which is used here on the repository owner's authority; no confidential, client, or personal data is involved, and direct identifiers are excluded from every response. Nothing here makes an approval, denial, pricing, underwriting, or automated decision, and no fraud-detection capability is claimed.

Deployment and cost boundary

ProviderRuntimeStateExposureVerifiedProduction claim
Azure Container AppsSingle Docker container on the Consumption plan, 1.0 vCPU / 2.0 GiB, scale-to-zero: FastAPI serves the built React application, the JSON API and the report artifacts from one process, reading a SQLite star schemaliveanonymous2026-08-10T15:19:17ZNo

Known limitations

  • The outcome is first-EMI default, not fraud and not long-horizon credit performance; the source has no fraud flag and no channel field.
  • A 21.7% first-EMI default rate implies the label mixes operational and mandate failures with genuine credit risk, and nothing in the source separates them.
  • Loss severity is an ultimate-loss assumption applied to a first-payment miss; the source has no recovery or roll-rate data.
  • Three months of 2018 disbursements, evaluated retrospectively, with October in festival season. Nothing here establishes forward performance.
  • The money figures are assumption-led. No band clears zero at the published assumptions, and in 13 of 27 published combinations none clears zero either.
  • The best band's net contribution is not statistically distinguishable from zero: a 95% interval of −₹0.65 to +₹0.49 cr.
  • Predicted probabilities are miscalibrated on the holdout by 3 to 5 points in every decile; the scores rank but do not price.
  • Under the conservative band 11 of 18 readable state codes fall below a four-fifths approval screen. The source has no protected attribute, so this is impact reporting, not a fair-lending test.
  • There is no status-quo baseline to compare against, because no incumbent policy is known for this dataset.

Scalability roadmap

  1. The store is the binding constraint, not the model. SQLite serves 233,154 rows from a single file well past this demo's needs; a genuine origination volume with concurrent writers would move it to Postgres, and the shared parameterized SQL layer is the seam that makes that a swap rather than a rewrite.
  2. Scoring is offline and batch. Real-time decisioning would need the calibrated model behind a service with its own latency budget and a feature store, neither of which exists here and neither of which this retrospective evidence would justify building yet.
  3. The economics are recomputed in the browser from three per-band totals because the risk cut-off does not move with the assumptions. Any change that lets an analyst move the cut-off itself breaks that shortcut and needs the scoring path back on the server.
  4. Nothing is monitored. Before this carried an operational decision it would need alerting, a scheduled reproducibility check against the published hash, and drift monitoring on the score distribution.

All projects