2026

AI systems now measure living bodies and act on what they find. No standard says when to trust them.

FEAIS is a research foundation. Our subject is a chain that existing safety work tends to split into pieces: a reading taken from air or tissue, a model's interpretation of it, and a physical action that follows. Error compounds along the chain, and the last step is rarely reversible by default.

The central question

What makes a system safe when it measures a body and then acts on it?

Calibration drifts. Histology samples a fraction of a percent of the tissue. Plasma markers stand in for processes nobody can watch directly. None of this is news to the people who run these instruments, and practice has built habits around all of it.

What is new is putting a model in the loop that has access to none of those habits. It sees a number, not the six months of humidity that produced the number. Once its output drives an air handler or a clinical recommendation, the uncertainty dropped at the front end comes back as a physical decision.

Research areas

Three research areas

The numbering follows the direction errors travel. Instrument error propagates into inference, and inference licenses action. Pieces of this are well covered already: metrology at the front end, model auditing in the middle, safety engineering for robots and device software at the far end. What nobody is doing is testing the chain end to end, which is where the compounding happens.

AIR, SURFACE TISSUE, ORGANISM WORLD 01 02 03 Sense Infer Act instrument model actuator a value a class calibration error label noise discarded at the interface an irreversible action taken on a point estimate
Figure 1  Each stage quantifies its own error, and the interface between stages does not carry it. The instrument knows its calibration uncertainty and the classifier knows its label noise, but what crosses each boundary is a point estimate. By stage three the system is acting on a number that has lost every qualification attached to it. Model confidence at stage two is not a substitute, because it is computed with respect to the training distribution rather than the instrument.
Total ion current
Sample VOC reference mixture
Column DB-5ms, 30 m
Run 21.0 min

Stage one produces something like this. Peak position identifies a compound, peak area quantifies it, and both carry error that stages two and three never see. Illustrative trace, indicative retention times.

01 Layer
Air, surface

Biological sensing

Olfaction, volatile organic compound detection, exposure monitoring and the chemical interface between a device and skin or mucosa.

  • Q.1Can an array separate a hazard from ordinary variation when both sit in the low parts per billion?
  • Q.2After six months of humidity cross-sensitivity and baseline drift, what is the instrument actually reporting?
  • Q.3Who can spoof the front end, and how cheaply?
02 Layer
Tissue, organism

Biological interpretation

Computational pathology, tissue imaging, biomarker panels and inference about disease or exposure from evidence that was never complete.

  • Q.4What happens downstream from a call of infection or neurodegeneration that turns out to be wrong?
  • Q.5How do you audit a classifier whose training labels were uncertain in the first place?
  • Q.6If a prediction rests on a proxy, can the system be made to name the proxy?
03 Layer
World

Embodied action

Robotics, self-driving laboratories, medical device software and anything else where an output moves matter.

  • Q.7How should a system behave when its confidence is low and the action is irreversible?
  • Q.8Which actions should be made reversible in hardware rather than in policy?
  • Q.9When a person is escalated to, what is in front of them, and how long do they have?

State of the art

These systems already work well enough to be deployed

Every row below is a published result rather than a projection of one. The last column is what that system would need before anyone lets it act without a person checking. None of those things exist yet, which is the work.

StageSystemReported resultSourceWhat is needed
01 Breath analysis across 17 disease classes A nanomaterial sensor array reached 86 per cent accuracy on 1,404 subjects, and gas chromatography identified 13 volatile compounds carrying the discrimination. Nakhleh 20171 A drift specification. The array was accurate on the day it was calibrated and nothing states how long that holds.
01 A learned map of odour percepts On 400 prospectively collected odorants the model matched the trained panel mean more closely than the median human panellist did. Lee 20232 A statement of what the map cannot resolve. Panel agreement on single odorants says little about a mixture in a room.
01 Breath screening for acute infection A multiplexed nanomaterial array separated COVID-19 cases from controls in exhaled breath during the pandemic. Shan 20203 A false positive budget agreed before deployment, because the response to an alarm is isolation of a person.
02 Whole-slide diagnosis at clinical grade Weakly supervised training on 44,732 slides from 15,187 patients gave areas under the curve above 0.98, enough to set aside 65 to 75 per cent of slides at full sensitivity. Campanella 20194 A check that the model reads the tissue rather than the hospital that submitted the slide.
02 Olfactory loss as a prodromal marker Impaired smell precedes motor onset in Parkinson's disease by years and is present in most early cases, which is what makes the nose worth instrumenting at all. Doty 20125 A rule for who is told. A marker that runs years ahead of treatment is a prediction somebody has to live with.
03 Autonomous chemistry A free-roaming robot ran 688 experiments over eight days across a ten-variable space and found photocatalyst mixtures six times more active than the starting formulations. Burger 20206 A stop that acts on the physical state of the bench and not only on the queue of planned experiments.
03 Population-scale biological monitoring A design for pathogen-agnostic metagenomic sequencing of wastewater, meant to catch exponential growth in an organism nobody has characterised yet. NAO Consortium 20217 A route from detection to a decision, with a named authority and a deadline attached to it.

Why olfaction

Olfaction is where we start

Smell is the one sensory system that already contains the whole problem. Molecules arrive, receptors encode them, the encoding is assigned a meaning, and an animal changes what it is doing. Each step has a counterpart in an embodied machine, so olfaction is somewhere to run experiments rather than a metaphor to argue about.

Olfactory system Embodied AI system
01 Encounters airborne molecules Physical signal
02 Converts them into a sensory representation Model of the world
03 Assigns meaning to that representation Decision
04 Produces an action Physical consequence

There is a practical reason too. Volatiles are one of the few routes by which information leaves an organism and enters shared air, which is what makes them attractive for detecting pathogens, decomposition and industrial exposure, and what makes the detectors worth attacking.

Biology also sets a benchmark that engineering has not met. Reported human detection thresholds for geosmin in water fall in the region of 4 to 15 nanograms per litre, which is single-digit parts per trillion. Arrays deployed in the field usually operate orders of magnitude above that, so asking whether they register the relevant compound at all is not a rhetorical question.

Digital pathology

Models that read air and tissue together

Air sampling records what an organism came into contact with. Histology and plasma record what followed. Groups are already building models that draw on both, and the audit problem is that the evidence enters at four depths, in units with no common denominator, carrying error structures that look nothing like each other.

Depth 01Air Airborne chemistryWhat was in the air MOX array, GC-MS for confirmation ppb to ppt
Depth 02Function Olfactory functionWhether the biological detector still works UPSIT, or Sniffin' Sticks TDI items correct
Depth 03Tissue Tissue morphologyWhat changed at the site Whole-slide H&E, 0.25 µm per pixel per field
Depth 04Systemic Molecular markersWhat changed away from the site Plasma pTau-217, GFAP, NfL pg / mL

Getting the classification right is hard and already well studied. The failure that worries us is quieter: a system reads exposure off a proxy, carries the correlation forward as though it were a mechanism, and hands a clinician an intervention. Acceptance criteria for that class of tool should exist before procurement decisions start to depend on them.

Submitting site Stain, scanner, fixation protocol Patient population, referral pattern Image features Reported outcome Tumour biology the path a clinician assumes is used the association a model can exploit instead
Figure 2  How a histology model scores well without learning biology. The submitting site sets how a slide is stained and scanned, and it also sets which patients walk through the door, so image appearance and recorded outcome share a common cause. Howard and colleagues found that site is recoverable from the image itself across more than 3,000 patients and six cancer subtypes, that it survives the colour normalisation normally used to remove it, and that it biases predictions of survival, mutation status and stage.8

Failure evidence

Published failures of the same kinds of system

None of the failures below are hypothetical. They fall into a few kinds. Some are a model learning the wrong variable. Some are an instrument ceasing to mean what it meant when it was calibrated. Some are an attacker going after the channel rather than the model. What they share is that internal validation does not reveal any of them.

01

The model learns the site rather than the disease

Across more than 3,000 patients and six cancer subtypes in TCGA, the submitting site is recoverable from the slide image itself. It survives the colour normalisation and augmentation used to remove it, and it biases predictions of survival, mutation status and stage. Site correlates with outcome through patient population, so a model can score well while never touching biology.

Howard et al., Nature Communications, 20218
02

Pathology foundation models still encode the hospital

A robustness index measuring whether biological signal or centre signal dominates the embedding space found that every pathology foundation model tested encodes the medical centre strongly, and that most cluster more tightly by centre than by cancer type. The errors are not random: they are confusions with other classes from the same centre.

de Jong et al., preprint, 202511
03

A model that passed validation on shortcuts failed in use

Chest radiograph models for COVID-19 were shown to draw on medically irrelevant correlates rather than pathology. The way the training data was assembled made this close to inevitable, and the deficit is invisible in internal validation. It appears as a performance collapse in a new hospital.

DeGrave et al., Nature Machine Intelligence, 202112
04

Gas sensors drift away from their calibration

Metal oxide arrays drift with humidity, ageing and exposure history. Over three years of controlled recording, drift degrades classifier performance measurably, and a later re-analysis showed the drift itself is informative enough to classify the analyte, meaning accuracy on the standard benchmark has been widely overstated.

Vergara et al., Sensors and Actuators B, 20129; Dennler et al., preprint, 202110
05

Sound can force a sensor to report a chosen value

Acoustic injection at a MEMS accelerometer's resonant frequency can make it deliver values of the attacker's choosing to the processor, which is a step beyond denial of service. No amount of model hardening helps here, because the model's input is exactly what the attacker specified. Any chemical or physical front end has the same exposure.

Trippel et al., IEEE EuroS&P, 201713
06

Invisible image edits flip a diagnosis and raise confidence

Imperceptible changes to a medical image flip a classification while raising confidence. The point made at the time was not the geometry of the attack but the incentive structure around it, since billing, triage and reimbursement all sit downstream of these outputs.

Finlayson et al., Science, 201914
07

A drug safety model run in reverse designed nerve agents

A toxicity model built to screen drug candidates for safety was run with the sign on its objective reversed. In under six hours it generated roughly 40,000 candidate agents, including VX and novel compounds predicted to be more toxic. A model trained to recognise a biological hazard inverts the same way, which is the part of dual use that monitoring systems own.

Urbina et al., Nature Machine Intelligence, 202215
calibration 1.0 0.8 0.6 NORMALISED RESPONSE TO A FIXED CONCENTRATION 0 6 12 18 24 30 36 MONTHS SINCE CALIBRATION accumulated offset
Figure 3  Schematic of the effect Vergara and colleagues recorded with a sixteen-sensor metal oxide array against six analytes over three years.9 The instrument does not fail; it slowly stops meaning what it meant. The sharper finding came later: Dennler and colleagues showed that residual drift in that same benchmark carries enough information to classify the gas on its own, so a substantial body of work using it has been reporting the drift signature rather than chemical discrimination.10 Curve illustrative, axis span real.

Failure mode register

AI that monitors biology, not AI that designs pathogens

Most of the AI and biology conversation is about design, and in particular whether a model lowers the barrier to building a pathogen. That work is serious and well funded, and we are not duplicating it. Far less attention goes to the tools meant to detect, interpret and contain, which have failure modes of their own. This is our working register of them.

IDFailure modeClassLayer
FM-01 Missed detectionThe agent was present and the instrument returned nothing unusual. False negative Air
FM-02 Over-triggeringNormal variation reads as an event, and the response costs more than the hazard would have. False positive Air, tissue
FM-03 Adversarial maskingA mixture is chosen to occupy the array's blind spot, or to saturate its response. Adversarial Air
FM-04 Contaminated autonomyA self-driving lab acts on bad data, then generates more data from the corrupted state. Cascade World
FM-05 Vector transferA robot carries biological material across a barrier that was there for a reason. Cascade World
FM-06 Uncontestable inferenceSomeone is assessed by a biological monitor they cannot inspect or appeal. Governance Body
PRIORITY QUADRANT hard to notice, hard to undo unnoticed obvious immediately DETECTABILITY AT THE MOMENT IT HAPPENS permanent undoable REVERSIBILITY OF THE CONSEQUENCE FM-03 adversarial masking FM-05 vector transfer FM-01 missed detection FM-04 contaminated autonomy FM-06 uncontestable inference FM-02 over-triggering
Figure 4  The register plotted on the two axes that decide how much engineering a failure mode is worth. Over-triggering is expensive but announces itself and can be walked back, which is why it tends to get the attention. The four modes in the shaded quadrant are quiet at the time and cannot be undone afterwards, and they are the ones a monitoring system has no established test for. Placements are our judgement rather than measured values.

How we differ

How this differs from AI safety and from biosecurity

AI safety concentrates on what a model returns into software. Biosecurity concentrates on what an actor could construct. Device regulation covers the last metre, but only once a product exists and a manufacturer is accountable for it. A model that reads a chemical or tissue signal and drives a machine falls between the three, and tends to be assessed as a device late, if at all.

Established field

AI safety

Channel
Text, code, software

Judges what the model returns.

The gap

FEAIS

Channel
Air, tissue, organism, machine

Judges what the return does physically.

Established field

Biosecurity

Channel
Organisms, agents

Judges what an actor could construct.

ComputationBiology

Contact

Get in touch

We are looking for people to work with and for one problem narrow enough to make real progress on. Disagreement with the framing on this page is as useful as agreement and easier to act on early.

Wanted

People we want to hear from

  • A regulatory or clinical background, ideally device software or human tissue
  • Research in chemical sensing, computational pathology or robotics
  • Governance experience on a scientific board
Wanted

A first project

One problem narrow enough to publish on inside a year: an instrument with a drift history, a slide set with known label noise, or a robot whose reversibility can be tested.

Contact

Write to us

Proposals and offers to collaborate reach the same address. So do objections.

contact@feais.org

References

Sources for everything claimed above

Figures on this page are drawn from these papers, or, where the caption says so, are our own schematics. No figure here reports a measurement of ours.

  1. Nakhleh MK, Amal H, Jeries R, et al. Diagnosis and classification of 17 diseases from 1404 subjects via pattern analysis of exhaled molecules. ACS Nano 11, 112 (2017). doi:10.1021/acsnano.6b04930
  2. Lee BK, Mayhew EJ, Sanchez-Lengeling B, et al. A principal odor map unifies diverse tasks in olfactory perception. Science (2023). doi:10.1126/science.ade4401
  3. Shan B, Broza YY, Li W, et al. Multiplexed nanomaterial-based sensor array for detection of COVID-19 in exhaled breath. ACS Nano (2020). doi:10.1021/acsnano.0c05657
  4. Campanella G, Hanna MG, Geneslaw L, et al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nature Medicine (2019). doi:10.1038/s41591-019-0508-1
  5. Doty RL. Olfactory dysfunction in Parkinson disease. Nature Reviews Neurology (2012). doi:10.1038/nrneurol.2012.80
  6. Burger B, Maffettone PM, Gusev VV, et al. A mobile robotic chemist. Nature 583, 237 (2020). doi:10.1038/s41586-020-2442-2
  7. The Nucleic Acid Observatory Consortium. A global nucleic acid observatory for biodefense and planetary health. arXiv (2021). arXiv:2108.02678
  8. Howard FM, Dolezal J, Kochanny S, et al. The impact of site-specific digital histology signatures on deep learning model accuracy and bias. Nature Communications (2021). doi:10.1038/s41467-021-24698-1
  9. Vergara A, Vembu S, Ayhan T, et al. Chemical gas sensor drift compensation using classifier ensembles. Sensors and Actuators B: Chemical 166, 320 (2012). doi:10.1016/j.snb.2012.01.074
  10. Dennler N, Rastogi S, Fonollosa J, van Schaik A, Schmuker M. Drift in a popular metal oxide sensor dataset reveals limitations for gas classification benchmarks. arXiv (2021). arXiv:2108.08793
  11. de Jong ED, Marcus E, Teuwen J. Current pathology foundation models are unrobust to medical center differences. arXiv (2025). arXiv:2501.18055
  12. DeGrave AJ, Janizek JD, Lee S-I. AI for radiographic COVID-19 detection selects shortcuts over signal. Nature Machine Intelligence (2021). doi:10.1038/s42256-021-00338-7
  13. Trippel T, Weisse O, Xu W, Honeyman P, Fu K. WALNUT: waging doubt on the integrity of MEMS accelerometers with acoustic injection attacks. IEEE European Symposium on Security and Privacy (2017). IEEE 7961948
  14. Finlayson SG, Bowers JD, Ito J, Zittrain JL, Beam AL, Kohane IS. Adversarial attacks on medical machine learning. Science 363, 1287 (2019). doi:10.1126/science.aaw4399
  15. Urbina F, Lentzos F, Invernizzi C, Ekins S. Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence (2022). doi:10.1038/s42256-022-00465-9