Ant Group’s AI lab inclusionAI unveiled Ling 3.0 Flash Sante on 4 September 2026 – a medicine-focused variant of its open-weight language model that is, for now, available only through programming interfaces.
What the model does
Sante builds on Ling 3.0 Flash, a mixture-of-experts model with 124 billion parameters, of which only around 5.1 billion are active per token. Its hybrid architecture pairs Kimi Delta Attention with Multi-Head Latent Attention and handles contexts of up to 262,144 tokens. The medical variant is tuned for clinical reasoning, evidence-based retrieval and long-horizon diagnostic tasks, while the vendor says it retains general strengths in logic and coding.
Open weights – with a caveat
The base model, Ling 3.0 Flash, ships under the permissive MIT licence on Hugging Face and scores roughly 38 on the Intelligence Index maintained by analysis firm Artificial Analysis – at a brisk pace of more than 300 tokens per second. For the Sante version, however, inclusionAI has not released the weights so far; it runs exclusively through providers’ APIs.
Handle the numbers with care
InclusionAI cites a DiagnosisArena-MCQ score of 83.8 for Sante – a vendor-run benchmark that no independent body has yet confirmed. The serving platforms also state plainly that the system is not a medical device and does not replace professional medical judgement. For clinics and developers in Germany, that leaves open how the model performs in practice, beyond the marketing claims.
Sources: Artificial Analysis · OpenRouter · Hugging Face



















