
TOXmod
AI Toxicity Prediction for Drug Discovery
TOXmod predicts molecular toxicity across multiple biological endpoints from a single SMILES string. Where most tools return a binary toxic / non-toxic label, TOXmod returns 18 quantitative, explainable and regulatory-aligned outputs per endpoint — calibrated uncertainty, applicability domain scoring, and atom-level structural guidance a chemist can act on directly. Below, the hERG cardiotoxicity endpoint is shown in full.
How It Is Built
The Problem We Solve
Every toxicity endpoint has its own failure story. Cardiotoxicity is the one that ends the most programmes — so it is the endpoint we walk through here.
Why hERG matters
- The hERG potassium channel repolarizes the cardiac action potential. Drugs that block it prolong the QT interval and can trigger fatal arrhythmia (Torsades de Pointes).
- Terfenadine, Cisapride and Grepafloxacin were all withdrawn from market due to hERG liability.
- Published literature attributes roughly 30% of drug withdrawals since 1990 to cardiac causes.
- FDA ICH S7B requires both in-vitro and in-silico hERG assessment for every IND filing.
- Industry estimates put the cost of a Phase III failure at $300M–$500M per programme.
Where existing tools stop short
- Binary toxic / non-toxic classifiers carry no quantitative potency information — a compound at 0.05 µM is treated the same as one at 8 µM.
- No uncertainty quantification, so predictions are presented to chemists as certainties.
- No applicability domain, meaning models get applied well outside the chemical space they are valid for.
- No atom-level explainability — a chemist learns a compound is flagged, but gets no direction for redesign.
- Desktop-bound licences that integrate poorly with modern automated screening pipelines.
What Comes Back
One SMILES string in. A structured report of 18 outputs across six functional groups out — quantitative, explainable and regulator-ready. Shown here for hERG.
IC₅₀ Regression
- hERG IC₅₀ as a continuous value in µM — every downstream metric derives from this number
- pIC₅₀, the universal unit across literature and ChEMBL, with a 95% confidence interval
Derived Safety Metrics
- Safety margin (IC₅₀ / Cmax) — the figure that appears in IND safety dossiers
- Redfern risk classification, expressed in ICH S7B regulatory vocabulary
- Molecular descriptors: MW, logP, TPSA, HBD and HBA
- A configurable pass / fail decision flag for batch screening pipelines
- A plain-text explanation summary for non-computational stakeholders
Potency Classification
- Four-class potency label: strong, moderate, weak or non-blocker
- Calibrated probabilities across all four classes, not just the winning label
- A prediction confidence score — high, medium or low
- A confidence-adjusted decision spanning five actionable outcomes
Uncertainty Quantification
- Epistemic uncertainty — where the model lacks training data in this region of chemical space
- Aleatoric uncertainty — the inherent noise in the underlying patch-clamp assay
- A combined 95% confidence interval on IC₅₀, honest about both sources
Explainability (XAI)
- Global SHAP feature importance, validating that predictions rest on known hERG pharmacophores
- Per-compound SHAP attribution, signed by whether each feature raises or lowers risk
- An atom-level 2D contribution map — red atoms increase risk, blue atoms are protective
Applicability Domain
- An applicability domain score against the training set, so out-of-domain compounds are flagged
- The nearest-neighbour training compound and its measured IC₅₀
- SMARTS-based structural alerts for known hERG toxicophores
Built For Three Roles
Medicinal Chemist
Rapid triage of lead compounds, with atom-level guidance on exactly which part of the structure to modify — no understanding of SHAP or machine learning required.
Computational Scientist
Quantitative IC₅₀ regression, calibrated uncertainty, ensemble agreement, SHAP values and applicability domain scores, all accessible over an API.
Regulatory Scientist
ICH S7B Redfern classification, OECD QSAR-compliant applicability domain reporting, and plain-language summaries ready for submission appendices.
Use Cases
Lead Optimization
Triage a series early and redirect chemistry effort before synthesis and assay budget is committed to a compound carrying cardiac liability.
IND-Enabling Safety Packages
Produce safety margins and Redfern classifications in the vocabulary regulators expect, with plain-prose summaries for submission appendices.
High-Throughput Screening
Score large compound libraries through the API and filter on a single decision column before anything reaches the bench.
Regulatory Alignment
Every output maps to at least one regulatory standard, so predictions arrive in the vocabulary submissions already use.
ICH S7B
Safety margin calculation and Redfern risk classification are implemented in the exact regulatory vocabulary used for hERG cardiac liability assessment in IND submissions.
OECD QSAR
All five OECD validation principles are addressed: defined endpoint, unambiguous algorithm, defined applicability domain, goodness-of-fit, and mechanistic interpretation.
EMA 2024
EMA guidance on in-silico evidence requires probabilistic predictions, SHAP-based explanations and explicit applicability domain assessment. TOXmod provides all three.
FDA CiPA
The Comprehensive In-vitro Proarrhythmia Assay initiative requires multi-channel cardiac assessment. TOXmod covers the primary hERG channel, with further channels on the roadmap.
Extending the Toxicity Panel
The featurization pipeline, uncertainty module, explainability layer and applicability domain scoring are reusable across endpoints. These are the ones we are extending to next.

Sarvayu
Digital Twins for Type 2 Diabetes Progression
Sarvayu is currently in development. Everything below describes the intended design of the platform, not a released product.
Sarvayu is a digital twin framework for Type 2 Diabetes built on Wasserstein GANs with gradient penalty. It generates individualized replicas of patients that simulate longitudinal disease trajectories, letting clinicians and researchers explore what-if scenarios for treatment optimization, progression forecasting and personalized intervention planning.
Planned Capabilities
Use Cases
Personalized Care Pathways
Support precision medicine in endocrinology by forecasting a patient’s metabolic profile and timing interventions to the window where they matter.
Complication Risk Stratification
Surface patients trending toward nephropathy or cardiovascular complications early enough for care to be escalated proactively.
Treatment Pathway Simulation
Model differential outcomes between therapy options — such as monotherapy versus combination therapy — before committing a patient to one.

Curexa
AI Platform for Clinical Trial Optimization
Curexa is currently in development. Everything below describes the intended design of the platform, not a released product.
Curexa applies multi-modal machine learning to clinical trial design and management — forecasting risk, optimizing resources and surfacing the patterns that manual protocol design misses. It is built on curated, AI-ready datasets drawn from public and commercial trial registries, fusing tabular records, free text, molecular graphs, ICD-10 codes and MeSH terms.
Planned Predictive Tasks
Use Cases
Protocol Design
Stress-test a protocol before it goes live — forecast duration, model dropout, and draft eligibility criteria against historical trial outcomes.
Trial Risk Monitoring
Anticipate serious adverse event and mortality signals early enough to adjust the trial, rather than reacting once it has already been halted.
Portfolio Planning
Compare approval likelihood and likely failure modes across candidate programmes to direct R&D investment where it can succeed.
See TOXmod run on your compounds
We work alongside R&D and data teams at pharmaceutical companies, biotechs, and healthcare institutions — integrating our models into the pipelines and governance frameworks you already run.
Enterprise engagements only — every partnership starts with a technical evaluation against your own targets and data.