New — Independent LLM model validation for regulated industries, and evaluation datasets available off the shelf. View offerings
Model validation · evaluation engineering · expert data

Statistical proof for AI systems that have to be right.

InterClin AI is an independent statistical practice for AI: model validation for regulated industries (U.S. bank model-risk guidance, FDA, EU AI Act), evaluation and gold-data engineering with measured results, and expert data services for frontier programs. Every deliverable is documented — inter-rater agreement (κ) with confidence intervals, adjudication logs, versioned errata — by statisticians with a minimum of 15 years of regulated-industry experience, including FDA/EMA submissions.

FIG.01 — A production LLM judge scored against expert labels. Raw agreement looks acceptable. The chance-corrected picture — and the 44% of true failures waved through — does not.
15+ years
Minimum regulated-industry experience · every statistician
κ ≥ 0.7
Reliability target, set a priori, every gold set
100%
Of deliverables ship with documentation
48 hrs
Fixed-fee proposal after scoping call
Selected engagements Artera AI IDEAYA Biosciences Bayer HealthCare
USE CASES Independent LLM & ML model validation (bank model risk) Judge reliability uplift with measured results Gold datasets & RL verifiers Vendor data acceptance testing FDA credibility evidence packages
Standards we work to US Interagency Model Risk Guidance (ex SR 11-7) FDA AI Credibility Framework EU AI Act — High-Risk FDA–EMA Good AI Practice HIPAA-aware workflows
Where we operate

Where we operate in the AI value chain.

FIG.02 — The AI value chain. Capital layers are closed; InterClin AI operates where expertise, measured results, and regulatory mandates create demand.
Who we serve

Solutions for every stage of AI evaluation and validation.

Mandated

Banks, Fintech & Insurers

Independent validation of ML and LLM models under the U.S. interagency model-risk framework (formerly SR 11-7 / OCC 2011-12) — conceptual soundness, outcomes testing, and validation reports your examiners expect.

How it works
Mandated

Health & Life-Science AI

FDA credibility evidence, analytical and clinical validation design, and fractional statistical leadership for regulated health AI.

How it works
Measured results

AI Product Teams

Judge and eval reliability engineering in production — we measure, remediate, and re-measure, with evidence your enterprise customers can review.

How it works
Expert data

AI Labs & Data Vendors

Expert-designed evaluation tasks, gold-standard solutions, reward rubrics, and independent acceptance testing for data deliveries.

How it works
Services

Fixed-fee engagements, scoped like a statistical analysis plan.

Fixed-fee engagements with clearly defined scope. Each project begins with a signed analysis plan defining its acceptance metrics, and ends with documentation your stakeholders, examiners, and regulators can review. Proposals within 48 hours of a scoping call.

AI Model Validation (Regulated)
Mandated · Model risk / FDA

Independent validation of ML and LLM models for banking, insurance, and health AI: conceptual-soundness review, performance and stability testing, LLM-specific evaluation (robustness, hallucination, prompt sensitivity, bias), effective challenge, and a validation report built to the U.S. interagency model-risk framework (successor to SR 11-7 / OCC 2011-12, revised 2026) or FDA’s AI credibility framework.

Every validation is led and signed by a senior statistician with a minimum of 15 years of regulated submission experience — with ongoing monitoring plans available as a managed service.

Fixed fee per model
Scoped on request

Ongoing monitoring
Retainer available
Judge Reliability Uplift
Most popular · 3 weeks

Your LLM-as-judge, measured like a diagnostic test — then fixed, then measured again. We establish baseline agreement and Cohen’s κ against expert gold labels, run remediation (rubric engineering, decomposed judges, threshold tuning), and re-measure so the improvement is documented, not asserted.

Deliverables include before/after error rates with confidence intervals, the remediated judge configuration, and a labeling protocol your team can rerun.

Fixed fee
3 weeks · scoped on request
Eval Systems & Gold Datasets
Build

Evaluation infrastructure with real statistical power: error analysis on production traces, a failure taxonomy built to saturation, expert-labeled gold sets with documented inter-rater reliability, adjudication protocols, and pass/fail gates with confidence intervals your CI can enforce.

Eval systems & builds
Fixed fee · scoped
Gold Data, Verifiers & Vendor Audits
For labs & data buyers

Expert task design, gold-standard solutions, and reward rubrics for biomedical and statistical reasoning — built for AI labs, human-data vendors, and RL-environment teams. Every deliverable ships with full quality documentation: dual-rater reliability, adjudication logs, and versioned errata.

Buying data from a vendor instead? We acceptance-test the delivery before you sign off — sampling plans, independent re-label, a pass/fail number with a confidence interval.

Verifier & gold-set packs
Scoped / licensed

Vendor acceptance audit
Fixed fee · 2 weeks
Fractional Statistical Lead
Managed service · Health AI

Standing statistical ownership for health-AI and regulated teams: evidence-package strategy, validation and study design, evaluation program leadership, and regulatory-interaction support — the accountable senior statistician your program is expected to have, at a fraction of a full-time hire.

Monthly retainer
Scoped on request

Also available: externally controlled trial design and RWE comparators for life-science programs, and expert statistical support for AI-related disputes. Scoped individually.

κ
The κ standard

Documented results, not assertions.

Every engagement defines its acceptance metrics in a signed analysis plan before work begins — reliability targets for validations, measured improvement for uplift work — and every deliverable reports against those pre-registered metrics with confidence intervals, adjudication logs, and versioned errata.

Method

A validation methodology built on clinical-trial statistics.

01

Error Analysis

Read real production traces first. Build the failure taxonomy from evidence, not intuition, until new problem types stop appearing.

02

Gold Standard & Sampling

Stratification, sample sizes powered for the comparisons you’ll actually make, dual-labeled subsets for reliability.

03

Testing & Validation

Score models and judges like diagnostic tests — agreement, κ, per-mode sensitivity and specificity, bias, robustness, calibration.

04

Remediation & Gates

Fix what the numbers indict, re-measure the fix, and set confidence-interval gates your CI or MRM monitoring can enforce.

05

Report & Handoff

A severity-graded findings register, effective-challenge documentation, and protocols your team keeps.

Research

Open research on evaluation quality.

The Gold Standard Files is our published research series: benchmark audits, judge-validation studies, and labeling-methodology analyses — released with full methods every two weeks.

Subscribe to receive each issue by email. Subscribe →

FILE_001 — We measured an LLM judge like a diagnostic test. [SHIPPING THIS WEEK]
FILE_002 — How noisy is a famous public benchmark? [IN ADJUDICATION]
FILE_003 — Consensus vs. adjudication: head-to-head. [SAMPLING]
FILE_004 — What 400 traces actually buys you. [QUEUED]
FILE_005 — Anatomy of a reward rubric. [QUEUED]
About InterClin AI

A minimum of fifteen years validating evidence where being wrong isn’t an option.

InterClin AI is an independent statistical practice. Every statistician on our engagements brings a minimum of 15 years in regulated industries — clinical trial design, statistical analysis plans, diagnostics validation, real-world evidence, and FDA/EMA submissions across oncology, neurology, and rare disease.

We apply the statistical methods developed for drug approvals and model risk management — experimental design, diagnostic accuracy, inter-rater reliability, independent effective challenge — to AI model validation, training data, and benchmarks.

Clients work directly with the senior statistician responsible for their engagement — no intermediary layers, no outsourced analysis.

Contact

Building or evaluating AI systems? We can help.

Tell us about your model, your evaluation stack, or your data pipeline. We’ll scope the right engagement on a 20-minute call and deliver a fixed-fee proposal within 48 hours.

Request a proposal
Careers

Help build the statistical standard for AI.

InterClin AI is hiring its founding commercial and operations team — five roles where you own the function end to end and build the systems the firm will run on. Remote-first (US), New York preferred; full-time or contract-to-hire by mutual fit. Sales roles add uncapped commission. Click a role to expand.

Head of Growth Marketing (Founding)Marketing · Full-time or contract-to-hire · Remote (US)

Own the entire demand engine for a technical services brand: positioning, site conversion, our published research series, the newsletter, paid experiments, and the analytics that tell us what's working.

What you'll own

  • Growth strategy and channel mix across organic, content, communities, and targeted paid
  • The research-to-demand engine: packaging technical studies into campaigns that book calls
  • Site conversion, SEO, and measurement discipline

You bring

  • 5+ years marketing B2B technical services to enterprise or regulated buyers
  • Proof you've built pipeline from content at least once, with numbers
  • Comfort with statistics, AI evaluation, and model risk as subject matter
Apply for this role
Content & Demand Generation ManagerMarketing · Full-time or contract-to-hire · Remote (US)

Turn our technical work into the most credible content in the AI-evaluation niche: research write-ups, case studies, the newsletter, and webinars with tooling partners.

What you'll own

  • The editorial calendar: research releases every two weeks plus supporting assets
  • Newsletter growth, lifecycle email, and partner co-marketing logistics
  • Case-study production from client engagements, approvals handled properly

You bring

  • 3+ years in content or demand generation for technical audiences
  • Writing samples that explain quantitative material precisely
  • Operational discipline: calendars kept, assets shipped, metrics logged
Apply for this role
Founding Account Executive — Regulated & Enterprise AISales · Full-time · Remote (US) · Base + uncapped commission

Sell independent model validation, evaluation engineering, and expert-data services to banks, health-AI companies, AI product teams, and data vendors — consultative, technical, fixed-fee sales with a founder-statistician beside you on every serious call.

What you'll own

  • Full-cycle sales from qualified conversation to signed statement of work
  • The MRM-boutique and direct-to-bank channels; health-AI and vendor accounts
  • Proposal assembly with the founder; an honest forecast

You bring

  • 4+ years selling professional services or compliance products to enterprise or regulated buyers
  • Evidence of closing five- and six-figure engagements on value, not discounts
  • Fluency discussing validation, evaluation, or model risk with technical stakeholders
Apply for this role
Sales Development RepresentativeSales · Full-time · Remote (US) · Base + commission

Open doors with precision: researched, personally relevant outbound to model-risk leaders, health-AI executives, AI engineering teams, and data-vendor operators — booking qualified conversations for the founder and AE.

What you'll own

  • Target research and list-building across our named segments — no purchased lists, ever
  • Outbound sequences by email and LinkedIn, logged with full tracker discipline
  • Qualification calls, clean handoffs, and weekly feedback that sharpens the message

You bring

  • 1–3 years of outbound SDR/BDR work with a documented meetings-booked record
  • Writing that sounds like a smart human, not a sequence tool
  • Resilience and genuine curiosity about technical topics
Apply for this role
Operations & Client Delivery ManagerOperations · Full-time or contract-to-hire · Remote (US)

Run the machine so the statisticians can do statistics: contracts, scheduling, invoicing, subcontractor onboarding against our 15-year experience standard, data-handling compliance, and the delivery calendar across concurrent engagements.

What you'll own

  • Engagement operations end to end: MSAs/SOWs, kickoffs, milestones, invoicing, closeout
  • Vendor, subcontractor, and insurance administration; the conflict-check workflow
  • Data-handling compliance and the internal systems that keep quality boring

You bring

  • 4+ years in operations or client delivery at a professional-services firm
  • Experience with regulated or confidential client environments preferred
  • A visible love of checklists, version control, and things being where they should be
Apply for this role

InterClin AI is an equal-opportunity employer.