Deployed work

Data products you can open and inspect.

Use one broad domain and one broad role filter. Each project includes the problem, focus, implementation, evidence, source, and live product.

14 live projects

Home
Live projects

14 projects

Every item opens to the decision, implementation, evidence, and live product.

01Deployed technical prototype

Legal & knowledge

Legal Discovery Graph

Evidence-linked investigation across documents and entities

FocusAI engineer / Data engineer

ResultConnects documents, people, events, and citations in one inspectable investigation workflow.

PythonFlaskLangChainsentence-transformersONNX RuntimePostgreSQL + pgvector
02Deployed technical prototype

Legal & knowledge

Legal Document RAG

Grounded research over a controlled public corpus

FocusAI engineer / Data engineer

ResultTurns public legal documents into a searchable corpus with evidence-linked answers and an explicit refusal path.

PythonFlaskAzure Document IntelligenceAzure OpenAIAzure AI SearchAzure Blob Storage
03Working technical prototype

Finance & risk

Guarded Text-to-SQL

A review-first interface for natural-language analytics

FocusApplied AI engineer / Data engineer

ResultTurns generated SQL into a reviewable proposal and prevents direct model-to-database execution.

PythonFastAPIDuckDBSQLGlotAzure OpenAIMicrosoft Entra ID
04Technical project

Finance & risk

Automobile-Loan First-EMI Default Strategy Portfolio

A retrospective credit-policy platform over 233,154 Indian vehicle loans: portfolio and vintage reporting, score-decile and segment risk analytics, a policy workbench with editable economics, a per-loan inspector, and post-deployment monitoring.

FocusData Analyst

ResultA credit-policy analyst can test where to draw a first-EMI risk line, see the confidence interval around the answer and who it declines, and take a recommendation or a refusal to governance. On the published assumptions the honest output is a refusal: no evaluated band clears zero.

PythonJavaScriptscikit-learnSQLiteFastAPIReact
05Technical project

Finance & risk / Operations

Credit Risk Model Validation & Review-Capacity Lab

A retrospective validation workstation over 30,000 licensed academic credit records, built to compare a non-demographic model with simple references, quantify fixed review workloads, and inspect score-ranked sensitivity without making a lending decision.

FocusAnalytics / Model Validation

ResultThe selected model clears prevalence/random, repayment-delay, and logistic references on repeated paired development evidence, but ties calibrated Extra Trees because the PR-AUC advantage does not clear the prespecified practical margin. At 10% holdout capacity, 600 historical rows contain 431 observed defaults with 71.8% precision and 3.25× lift, each reported with uncertainty.

Pythonpandasscikit-learnPyArrowTypeScriptReact
06Technical project

Finance & risk / Operations

Application Fraud Strategy Portfolio

An application-fraud strategy desk over 1,000,000 loan applications: temporal model evaluation against promotion gates fixed before any result was seen, a capacity-bounded review queue wired to a PostgreSQL fraud schema, monthly KPI and vendor-performance reporting aggregated in SQL, and a daily suspect-application report from linking analysis.

FocusFraud Strategy Analyst

ResultA fraud strategy analyst can compare screening approaches at a fixed review capacity, see what each buys and costs, and take a recommendation or a refusal to governance. On the pre-agreed checks the honest output is a refusal: the proposed model catches 472 more fraud attempts while holding up 472 fewer good customers at identical cost, and is still not promoted, because its calibration and population stability fail checks written before the result was known.

PythonSQLPostgreSQLCatBoostscikit-learnPower BI
07Technical project

Healthcare

Synthetic Sepsis Risk Evaluation Lab

A synthetic-only model evaluation lab that compares a prevalence baseline with a calibrated challenger, exposes calibration and threshold tradeoffs, and keeps patient data outside the public system.

FocusData Scientist

ResultOn the fixed synthetic holdout, the calibrated Extra Trees challenger improves AUROC from 0.5000 to 0.9796 and Brier score from 0.0475 to 0.0228. The dashboard turns those model results into an inspectable threshold tradeoff without presenting them as clinical evidence.

Pythonscikit-learnFlaskDockerGitHub ActionsMicrosoft Azure
08Technical project

Healthcare / Operations

GP Access Planner

A public-data planning product that forecasts recorded general-practice appointments across England while keeping observed access signals and hypothetical capacity explicitly separate.

FocusSenior Healthcare Analyst

ResultDelivered a source-traceable 7, 14, and 28-day planning surface for 104 sub-ICBs, backed by 32.9 million validated source rows, rolling-origin evaluation, immutable releases, and a live edge API.

PythonPostgreSQLdbtscikit-learnLightGBMCatBoost
09Technical project

Customer & growth / Operations

Support Demand and Resolution Decision Briefing

A live planning workspace that turns licensed, masked helpdesk records into transparent weekday demand and priority resolution baselines without pretending to automate staffing decisions.

FocusAnalytics Engineering

ResultSupport leaders can frame one staffing-meeting question from two independent historical baselines, inspect the latest observed demand window, and make missing live inputs visible before discussing coverage.

PythonFlaskPostgreSQLDocker ComposeAWS LightsailCaddy
10Technical project

Finance & risk

Freddie Mac CRT Disclosure Analytics

A live collateral-surveillance workbench that connects portfolio delinquency movement to the complete masked Freddie Mac CRT loan-period disclosure record.

FocusAnalytics Engineer

ResultDelivered bounded public investigation across 20,439,666 masked loan-period rows, exact portfolio rate-and-mix attribution, private object storage, authenticated data retrieval, and a zero-incremental-cost production release.

PythonDuckDBParquetJavaScriptVercel FunctionsCloudflare Workers
11Technical project

Finance & risk

Payments Fraud Risk Data Platform

A governed fraud-risk validation register that publishes all 1,852,394 allowlisted simulated events for bounded analytical queries while keeping identity-like fields, model scores, and payment actions out of the public system.

FocusData Engineering

ResultRetain the ordinary logistic baseline for the measured 1% review queue. It reaches 0.160 PR-AUC and 51.3% recall while the class-weighted challenger performs worse on ranking, recall, and probability error.

PythonPostgreSQLscikit-learnSupabaseFastAPINext.js
12Technical project

Customer & growth

Subscriber Retention Intelligence

A governed retention analysis system over 442,211,685 accepted membership, payment, listening, and churn-label rows, with cutoff-safe cohorts, a calibrated repeat-subscriber model, and an assumption-bound intervention planner.

FocusAnalytics Engineer

ResultA subscription lifecycle analyst can locate material renewal and engagement changes, compare cohorts and segments on governed definitions, and test a finite contact program without confusing historical risk with causal treatment response.

PythonDuckDBdbt CoreParquetscikit-learnFastAPI
13Technical project

Finance & risk / Operations

Financial Payments Fraud Decision Workbench

A fraud strategy workbench built from the full 284,807-row MLG-ULB benchmark, with chronological evaluation, calibrated scoring, and a capacity-bounded review queue.

FocusData Scientist

ResultA fraud operations lead can compare threshold and staffing choices, see captured and missed fraud, inspect anonymized cases, and identify where extra review capacity stops paying back.

Pythonpandasscikit-learnXGBoostDashFlask
14Technical project

Finance & risk

Freddie Mac MBS Disclosure Intelligence

A public MBS disclosure decision product with exact source reconciliation, 38 released metric contracts, evidence-backed investigations, and a complete digest-verified data release.

FocusData Engineer

ResultDelivered a public trust-to-investigation workflow over 693,640,933 physical disclosure records, 264,922,553 loan-period facts, 9,240,038 security-period facts, and all 167 approved source and derived release artifacts.

PythonSQLitegzip CSVJavaScriptGitHub ActionsGitHub Pages