Kaimei
KAIMEIVietnam Japan AI Hackathon 2026 · Idea Proposal

01Project Snapshot

Agentic AI for Sustainable Goals

Health · Red Biotech (Drug Discovery)

KaimeiA computational lab for every biomedical scientist.

An AI agent that plans and runs early drug discovery with the scientist, on one canvas, with every result traced to its own steps.

SDG 3.4Non-communicable diseases

SDG 3.bResearch for new medicines

SDG 9.5Research capacity

02Problem & Impact

Bringing a new medicine to patients takes over a decade and more than a billion dollars, and nine in ten candidates that enter clinical trials still fail, most for lack of efficacy or for toxicity [1]. These failures surface late but are set early: in the first four stages, from target to preclinical, teams choose the target and the molecules on evidence still gathered and joined by hand. Kaimei focuses on these four stages: a sounder choice upstream means fewer failures downstream, an earlier start to development and, above all, safer medicines for patients.

Who Kaimei serves: Universities & institutes (researchers), biotech startups & pharma (bioinformaticians)

10-15+years from discovery to approval
$1-2B+per approved drug
90%fail in clinical trials
StageCandidatesCostYears
Target Validation-3%1.5
Compound Screening>10,0006%1.5
Lead Optimi­zation~25017%1.5
Pre­clinical10-207%1
Phase I~615%1.5
Phase II~421%2.5
Phase III~226%2.5
Approval15%1.5
Our focus in VJ AI Hackathon 2026
Discovered drug candidatesDrugs brought to marketDrug development

Evidence for one target is scattered across many tools and joined by hand.

One research question sends scientists through many tools: databases, analysis software, scripts and spreadsheets. Each tool keeps its own format, so data is copied, converted and matched by hand, costing hours and letting errors slip in.

Impact Every source and heavy model in one session, every step replayable.

Scattered across tools

UniProt FASTAChEMBL TSVPubChem JSONDocking SDFScript PythonExcel by hand
You

Agentic AI is spreading in drug discovery, yet the risk that agents leak lab data, invent results or act unasked gets little attention.

Discovery data is a lab's intellectual property. An agent with access can expose the data, obey an instruction planted in a document [2], report a value no run produced and leave no record of why.

Impact Every step and approval on record, in limits a person sets.

Inputs

Risks

Lab dataDocumentsModel runsYour requestData leakedPlanted instructionInvented valueNo record of why
AI agent
Project Snapshot · Problem & Impact1 / 3
KAIMEIVietnam Japan AI Hackathon 2026 · Idea Proposal

03Solution & Innovation

Six layers, one system. A scientist sets the goal and describes the work, approves the agent's plan, reads the report and checks the workflow pipeline on the canvas. Where today a scientist joins many tools by hand, Kaimei puts public databases and the lab's own files in one session and runs heavy biology models, such as structure prediction and docking, on cloud GPUs, inside the lab's secure environment. Each number traces back to a real run.

06 · What Kaimei shows

Canvas & Nodes

Every agent step becomes a node the team can read, change and rerun.

05 · What Kaimei does

Agentic AI

20+ specialist skills, every report checked against tool output.

04 · What Kaimei knows

Context & Knowledge

Knowledge from your files and papers, memory across sessions.

03 · What Kaimei runs

Biology & Small Molecule Tools

15+ heavy models ready as nodes: no install, no own GPU.

02 · What Kaimei reads

Data & Integrations

Public databases and lab files, read at the source and joined for you.

01 · Where Kaimei lives

Your Secure Environment

A tenant per organization, keys kept from the agent and its code.

04Agentic AI

Early discovery has no fixed recipe: every question needs its own sources, models and steps. The Kaimei agent reads the session, writes a plan a person approves, hands the parts to sub-agents that work side by side, reads every result and merges them into one sourced report.

Human Scientist1(Question, budget)
Orchestrator Agent3(Plans, splits the work)
Skills Hub(20+ specialist skills)
Memory2(Knowledge, context)
Agent Loop(Until done)
Target Agent(Finds, ranks targets)
Protein Agent(Folds, designs proteins)
Chemistry Agent(Docks, scores hits)
Literature Agent(Researches, cites papers)
Tool Hub(Biology & molecule tools, data, uploads)
Canvas Pipeline4(Builds workflows from nodes)
Review Gate(Checks sources)
Report5(Findings & reasoning)
  1. 1Goal
  2. 2Context / Data
  3. 3Decision
  4. 4Action
  5. 5Result

A person decides

  • The goal and the budget
  • The plan, before any run
  • The scientific call
  • Stop, at any time

Kaimei does alone

  • Reads the session each step
  • Asks, plans, splits the work
  • Builds the canvas workflow
  • Keeps going unattended

Code checks, not a model

  • Misuse screen on every call
  • Injection guard, cost ceiling
  • Tables checked against runs
  • Citations checked at source
Solution & Innovation · Agentic AI2 / 3
KAIMEIVietnam Japan AI Hackathon 2026 · Idea Proposal

05Responsible AI

Human Oversight

Risk: no one checks the work

A person approves each plan and can Stop any time, every step logged

AI Autonomy

Risk: acts or spends unasked

Never approves its own plan, acts within fixed token and cost limits

Hallucination

Risk: an invented value

Each value traced to its run, the check tested on planted errors

Misuse

Risk: a harmful molecule

Controlled-substance screen and refusal rules on every molecule

Privacy

Risk: private data exposed

Needs no patient data, each lab in its own tenant

Data Security

Risk: a planted instruction

Encrypted keys, keyless sandbox, injection guard on inputs

Intellectual Property

Risk: lab data reused

Kaimei never trains on lab data, sources cited with licences

Bias

Risk: known targets favoured

Accuracy compared on well-studied and rarely studied targets

06MVP & Feasibility

MVP core features

01 · Canvas & Nodes

Every experiment,
on one canvas.

A Kaimei canvas lays every experiment out as connected nodes, so you can follow inputs, model runs, structures, results and next steps without losing track.

Technical: Next.js, React, TypeScript, Mol* Viewer

02 · Models & Tools

Heavy models,
ready to run.

Kaimei sets up the latest biology models and chemistry tools as nodes, so you can run compute-heavy science without installing software or owning a GPU.

Technical: Python, FastAPI, NVIDIA GPU, API Integration

03 · Agents & Skills

Whole research tasks,
planned with you.

The agent picks a skill and runs the steps you approve, so you can rank targets, profile a compound or map how a ligand binds without writing the code yourself.

Technical: Python, Docker, LLM, BioLLM Model

Product Roadmap

H2 2026H1 2027H2 2027H1 2028
Business developmentSeek investors in the field and shape the product for a global marketPilots with labs in Vietnam and Japan, sales in Japan: Takeda, Chugai
ProductAgentic AI for every drug phase, from omics to approvalLab in the loop: wet-lab results feed the agent's next round
Cert/Regulation2G3M, SOC Type 2HIPAA, ISO/IEC 27001

Members

  • Nghia Pham

    Leader, PO, Engineer

    Product direction, build solution, pitch

  • Hue Lam

    BA, Tester

    Business requirements, design and UAT

Responsible AI · MVP & Feasibility3 / 3
KAIMEIAppendix A · Background

Appendix A · Background: how a medicine is found

Supporting material for readers outside biology. The graded sections are on pages 1 to 3.

Target (EGFR)PocketMedicine (erlotinib)

A target and its key

A disease often runs through one protein, its target. A medicine fits a pocket on that protein like a key in a lock. EGFR is stuck on in some lung cancers, and erlotinib fits its pocket (structure 1M17 [3]).

Three kinds of medicine

Small molecule

A small lab-made chemical, as most pills are: erlotinib.

In Kaimei: docking, ADMET, generation

Antibody

A large protein that binds the outside: cetuximab, 1YY9 [4].

In Kaimei: numbering, interface scoring

Designed protein binder

A new protein designed to bind a chosen site.

In Kaimei: RFdiffusion, ProteinMPNN

The example session, phase by phase

  1. 1 · Early discovery

    Target ranking. Start from the disease and see which targets its evidence ranks first.

    Person approves, agent ranks [5]

  2. 2 · Target validation

    Target assessment. Weigh the target first: its drugs, its risks, its structures.

    Agent reads safety data, fetches 1M17

  3. 3 · Hit to lead

    Library screen. Screen a ChEMBL library, then check the report against the tools' output.

    Agent screens [6], report sent back

  4. 4 · Hit to lead

    Approved antibody. The same target as an antibody: an approved complex, its interface.

    Agent places nodes, person presses Run

  5. 5 · Lead optimization

    Hits to leads. Dock the hits into a real structure, then grow analogues of the leads.

    Person removes two ADMET gates, presses Run

  6. 6 · Clinical

    Trial landscape. See where the drugs for this target stand in trials today.

    Agent searches ClinicalTrials.gov [7]

An illustrative session. The target's evidence, the structures, the reference drugs, the trial counts and the tools are real, as of September 2026. The compounds and their scores are examples.

Glossary

Target.
The protein a medicine acts on.
Hit, lead.
A hit is active. A lead is worth optimizing.
SMILES.
A molecule written as one line of text.
PDB id.
The code of a solved 3D structure.
Docking.
Predicting how a molecule sits in a pocket.
ADMET.
Absorption, distribution, metabolism, excretion, toxicity.
hERG, DILI.
Safety flags: heart rhythm, liver injury.
pChEMBL.
ChEMBL's potency scale, higher is stronger.
Node.
One step on the canvas.
Skill.
A written method for one whole task.
Plan.
Kaimei waits for a person to approve each plan.
Sandbox.
A sealed workspace for the agent's code.
How a medicine is found · The example session · GlossaryAppendix A
KAIMEIAppendix B · Tools, skills and sources

Appendix B · Tools, skills and sources

Supporting material: what the agent works with, and the sources the text cites.

15+ tools in one workspace

StructureBoltz-2 · ESMFold · OpenFold2 · AlphaFold DB
DockingBoltz-2 · BoltzMol-1 · DiffDock
DesignBoltzProt-1 · BoltzMol-1 · RFdiffusion · ProteinMPNN · MolMIM
PropertiesBoltz-2 · Boltz ADME · PubChem · ChEMBL · RDKit · PRODIGY · ADMET-AI
DatabasesAlphaFold DB · RCSB PDB · UniProt · PubChem · ChEMBL

22 skills the agent follows

DiseasesRank targets for a disease · Map the clinical trial landscape
Genes and variantsCheck genetic constraint · Tier a gene's clinical variants · Count alterations across cancer types
TargetsAssess druggability by modality · Decide whether to inhibit or activate · Find where a target is expressed
CompoundsProfile potency and selectivity · Check drug-likeness · Predict off-target liabilities · Train ADME models on your assays · Model potency and find new scaffolds
Structure and designMap how a ligand binds · Generate new molecules · Predict a protein's structure · Design binders against a target · Humanize and assess an antibody · Screen a library by docking
Literature and labReview the literature on a question · Review the preclinical evidence · Lay out a microplate experiment

References

  1. [1] Sun D, Gao W, Hu H, Zhou S. Why 90% of clinical drug development fails and how to improve it? Acta Pharmaceutica Sinica B. 2022;12(7):3049-3062. doi:10.1016/j.apsb.2022.02.002
  2. [2] Huynh DL, Seal S, Chelbi M, et al. AI agents in drug discovery: applications and case studies. Drug Discovery Today. 2026;31(3):104650. doi:10.1016/j.drudis.2026.104650
  3. [3] Stamos J, Sliwkowski MX, Eigenbrot C. Structure of the epidermal growth factor receptor kinase domain alone and in complex with a 4-anilinoquinazoline inhibitor. Journal of Biological Chemistry. 2002;277(48):46265-46272. doi:10.1074/jbc.M207135200. PDB entry 1M17.
  4. [4] Li S, Schmitz KR, Jeffrey PD, Wiltzius JJW, Kussie P, Ferguson KM. Structural basis for inhibition of the epidermal growth factor receptor by cetuximab. Cancer Cell. 2005;7(4):301-311. doi:10.1016/j.ccr.2005.03.003. PDB entry 1YY9.
  5. [5] Ochoa D, et al. The next-generation Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic Acids Research. 2023;51(D1):D1353-D1359. doi:10.1093/nar/gkac1046
  6. [6] Zdrazil B, Felix E, Hunter F, et al. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Research. 2024;52(D1):D1180-D1192. doi:10.1093/nar/gkad1004
  7. [7] ClinicalTrials.gov, U.S. National Library of Medicine. clinicaltrials.gov, searched September 2026.
Tools · Skills · ReferencesAppendix B