- Ant's finance-tuned Ling model now has downloadable weights, following the August API announcement.
- The model card describes a 124B-parameter MoE with 5.1B activated parameters and a 256K context window.
- FinFIRST publishes 123 test tasks and a documented protocol for checking financial source use.
From an announced release to available weights
Ant Group's finance-focused Ling model is now available as a downloadable checkpoint under inclusionAI, following the August announcement that weights would be released after API access. The current model card identifies Ling-3.0-flash-Fin as a finance-enhanced version of Ling-3.0-flash, developed through continued training with financial data and domain specialists.
Ant describes 124 billion total parameters, 5.1 billion activated per token and a 256,000-token context window. The checkpoint is released in BF16 with an MIT license and shares deployment architecture with the base model. SGLang and vLLM are among the documented serving routes.
The work the model targets
The model's stated focus is a connected research process: retrieve sources, reconcile evidence, calculate, build a financial model and prepare a reviewable report. Its card specifically discusses conflicting numbers across filings, different reporting periods, spreadsheet formulas and cross-sheet dependencies.
These are concrete sources of error in financial research. A correct number from the wrong reporting period can produce a misleading comparison. A spreadsheet that looks complete can still contain a broken formula or a hard-coded assumption where an editable dependency was expected. Ant is positioning the specialized model around these multi-step tasks rather than around general conversational knowledge alone.
How the public financial test is scored
FinFIRST's public test split contains 123 tasks in Chinese and English. The dataset was developed by Ant with professional support from CICC's investment-banking team. Its task records separate time constraints, source requirements, calculation types and the units or format required in the answer. That structure exposes why retrieving a plausible number is only one part of satisfying a financial question.
The benchmark reports fifteen model configurations evaluated in the same ReAct-style framework with the same search, page-reading and Python tools. GLM-5.1 is used as an automated judge. The authors checked that judge against fifty sampled instances reviewed by eight finance professionals, reporting Cohen's kappa of 0.816 at the individual-criterion level. This measures agreement on those labels; it is not the accuracy of the released Ling model on all financial work.
The distinction matters for comparison: an end-to-end score depends on the search results, tools, model and judge used in the experiment. The dataset's open rubrics allow researchers to inspect that measurement process alongside the model weights.
An accompanying test of source use
The released FinFIRST dataset gives a more tangible view of that objective. Its public examples contain questions, expected answers and detailed scoring criteria. One asks an agent to trace a pharmaceutical company's earnings forecast through revisions and its final annual report. The scoring separates identifying official disclosures from retrieving amounts and explaining why the figures changed.
Another example requests model API prices as of a specified date, with consistent currency and units. The significance is the time boundary and comparison method, not whether those historical prices remain current. Such tasks test whether an agent can respect an information cutoff and preserve the meaning of a number across several sources.
Ant reports evaluation across financial retrieval, spreadsheet and banking benchmarks. The release does not establish that every long research workflow is reliable; its own model card identifies complex multi-step work as an area needing further validation. The public weights and task-level dataset make the specialized release more inspectable than the original API announcement alone.
Sources & context
Go to the original material. Company claims remain attributed to their sources.
01Updates & corrections
— Updated the August announcement to reflect the subsequent availability of model weights and FinFIRST evaluation data.



