A single patient session can carry a diagnosis, a billing complaint, and an unrelated photo. Feed that noise to a vision-language model and it drifts or hallucinates. MANN-Engram filters it: distill the intent in the cloud, route the right evidence at the edge, pass only what is verified.
Cloud intelligence for language; edge intelligence for perception. Meeting in a shared latent space.
A messy narrative goes to Qwen-2.5-72B, which returns two purified fields: the Clinical Intent, and an image-recognition-style Visual Query — nouns and modalities only, no abstract symptoms.
▸The visual query and every candidate scan are embedded by a skew-Gaussian-optimized SiGLIP engine and scored by sigmoid-matching probability — all locally, on the edge.
▸Relative-margin thresholding keeps only candidates within a dynamic window of the best match. Irrelevant scans never reach the model’s context window.
This reproduces the router’s exact pruning rule. Drag each candidate’s match probability, and the routing tolerance — watch which scans survive.
A model is only as honest as the context you hand it. Prune irrelevant evidence and you remove the raw material of hallucination — and you do it at the edge, so the patient’s data never leaves the device to be filtered.
/execute_pipeline mirrors these inputs and metrics.