InterClin AI is an independent statistical practice for AI: model validation for regulated industries (U.S. bank model-risk guidance, FDA, EU AI Act), evaluation and gold-data engineering with measured results, and expert data services for frontier programs. Every deliverable is documented — inter-rater agreement (κ) with confidence intervals, adjudication logs, versioned errata — by statisticians with a minimum of 15 years of regulated-industry experience, including FDA/EMA submissions.
Independent validation of ML and LLM models under the U.S. interagency model-risk framework (formerly SR 11-7 / OCC 2011-12) — conceptual soundness, outcomes testing, and validation reports your examiners expect.
How it worksFDA credibility evidence, analytical and clinical validation design, and fractional statistical leadership for regulated health AI.
How it worksJudge and eval reliability engineering in production — we measure, remediate, and re-measure, with evidence your enterprise customers can review.
How it worksExpert-designed evaluation tasks, gold-standard solutions, reward rubrics, and independent acceptance testing for data deliveries.
How it worksFixed-fee engagements with clearly defined scope. Each project begins with a signed analysis plan defining its acceptance metrics, and ends with documentation your stakeholders, examiners, and regulators can review. Proposals within 48 hours of a scoping call.
Independent validation of ML and LLM models for banking, insurance, and health AI: conceptual-soundness review, performance and stability testing, LLM-specific evaluation (robustness, hallucination, prompt sensitivity, bias), effective challenge, and a validation report built to the U.S. interagency model-risk framework (successor to SR 11-7 / OCC 2011-12, revised 2026) or FDA’s AI credibility framework.
Every validation is led and signed by a senior statistician with a minimum of 15 years of regulated submission experience — with ongoing monitoring plans available as a managed service.
Your LLM-as-judge, measured like a diagnostic test — then fixed, then measured again. We establish baseline agreement and Cohen’s κ against expert gold labels, run remediation (rubric engineering, decomposed judges, threshold tuning), and re-measure so the improvement is documented, not asserted.
Deliverables include before/after error rates with confidence intervals, the remediated judge configuration, and a labeling protocol your team can rerun.
Evaluation infrastructure with real statistical power: error analysis on production traces, a failure taxonomy built to saturation, expert-labeled gold sets with documented inter-rater reliability, adjudication protocols, and pass/fail gates with confidence intervals your CI can enforce.
Expert task design, gold-standard solutions, and reward rubrics for biomedical and statistical reasoning — built for AI labs, human-data vendors, and RL-environment teams. Every deliverable ships with full quality documentation: dual-rater reliability, adjudication logs, and versioned errata.
Buying data from a vendor instead? We acceptance-test the delivery before you sign off — sampling plans, independent re-label, a pass/fail number with a confidence interval.
Standing statistical ownership for health-AI and regulated teams: evidence-package strategy, validation and study design, evaluation program leadership, and regulatory-interaction support — the accountable senior statistician your program is expected to have, at a fraction of a full-time hire.
Also available: externally controlled trial design and RWE comparators for life-science programs, and expert statistical support for AI-related disputes. Scoped individually.
Every engagement defines its acceptance metrics in a signed analysis plan before work begins — reliability targets for validations, measured improvement for uplift work — and every deliverable reports against those pre-registered metrics with confidence intervals, adjudication logs, and versioned errata.
Read real production traces first. Build the failure taxonomy from evidence, not intuition, until new problem types stop appearing.
Stratification, sample sizes powered for the comparisons you’ll actually make, dual-labeled subsets for reliability.
Score models and judges like diagnostic tests — agreement, κ, per-mode sensitivity and specificity, bias, robustness, calibration.
Fix what the numbers indict, re-measure the fix, and set confidence-interval gates your CI or MRM monitoring can enforce.
A severity-graded findings register, effective-challenge documentation, and protocols your team keeps.
The Gold Standard Files is our published research series: benchmark audits, judge-validation studies, and labeling-methodology analyses — released with full methods every two weeks.
Subscribe to receive each issue by email. Subscribe →
InterClin AI is an independent statistical practice. Every statistician on our engagements brings a minimum of 15 years in regulated industries — clinical trial design, statistical analysis plans, diagnostics validation, real-world evidence, and FDA/EMA submissions across oncology, neurology, and rare disease.
We apply the statistical methods developed for drug approvals and model risk management — experimental design, diagnostic accuracy, inter-rater reliability, independent effective challenge — to AI model validation, training data, and benchmarks.
Clients work directly with the senior statistician responsible for their engagement — no intermediary layers, no outsourced analysis.
Tell us about your model, your evaluation stack, or your data pipeline. We’ll scope the right engagement on a 20-minute call and deliver a fixed-fee proposal within 48 hours.
Request a proposalInterClin AI is hiring its founding commercial and operations team — five roles where you own the function end to end and build the systems the firm will run on. Remote-first (US), New York preferred; full-time or contract-to-hire by mutual fit. Sales roles add uncapped commission. Click a role to expand.
Own the entire demand engine for a technical services brand: positioning, site conversion, our published research series, the newsletter, paid experiments, and the analytics that tell us what's working.
Turn our technical work into the most credible content in the AI-evaluation niche: research write-ups, case studies, the newsletter, and webinars with tooling partners.
Sell independent model validation, evaluation engineering, and expert-data services to banks, health-AI companies, AI product teams, and data vendors — consultative, technical, fixed-fee sales with a founder-statistician beside you on every serious call.
Open doors with precision: researched, personally relevant outbound to model-risk leaders, health-AI executives, AI engineering teams, and data-vendor operators — booking qualified conversations for the founder and AE.
Run the machine so the statisticians can do statistics: contracts, scheduling, invoicing, subcontractor onboarding against our 15-year experience standard, data-handling compliance, and the delivery calendar across concurrent engagements.
InterClin AI is an equal-opportunity employer.