News

Ant Group Bailing open source Ling-3.0-flash model: 124B total parameters, 5.1B activation parameters

2 min read
Source: ithome.com
IT House reported on August 8 that Ling Model, a subsidiary of Ant Group, officially announced yesterday that the new generation of native hybrid inference model Ling-3.0-flash is now officially open source. Ling-3.0-flash adopts MoE architecture with 124B total parameters and 5.1B activation parameters. The official version simultaneously provides the basic version and multiple versions such as FP8, FP4, INT4, etc., providing three different implementation options: API call: For developers who want quick access and do not need to build their own inference services, Ling-3.0-flash can be called directly through the cloud API; Stand-alone privatization: For enterprises and teams whose data cannot go out of the domain, the MXFP4 and INT4 versions can complete end-to-end inference on a single NVIDIA DGX Spark; High-performance deployment: For high-performance services that are sensitive to single request latency, under the specified GPU test configuration, the average output rate exceeds 1100 tokens/s. For developers and product teams who want to quickly verify products and launch Agent applications, API is the access method with the lowest threshold. In the Artificial Analysis list, Ling-3.0-flash output speed reaches 353 tokens/s. In the AA Intelligence Index list, the weighted average call cost of each task of Ling-3.0-flash is about US$0.04 (current exchange rate is about 0.27 yuan), and the weighted average decoding time is about 1.4 minutes. It also enters the advantageous areas of "intelligence level - task cost" and "intelligence level - task time consumption". Officials stated that from 00:00 on August 7, 2026 to 24:00 on August 31, Ling-3.0-flash can enjoy a 25% discount in Ling Studio for a limited time; starting from 00:00 on September 1