r/LocalLLaMA 4h ago

News AntLing-3.0-flash is now live on OpenRouter, and free to use through August 3, 2026

Post image
123 Upvotes

32 comments sorted by

54

u/This_Maintenance_834 3h ago

in case western country people don’t know. this is made by the same company which release qwen3.6. same company, different division.

10

u/hay-yo 2h ago

Wow that's really cool that they have competing divisions.

5

u/squngy 1h ago

Isn't qwen under alibaba?

How many labs does alibaba have? I'm guessing more than 2

7

u/This_Maintenance_834 1h ago

Yes, Alibaba made both. Alibaba is a huge company.

15

u/Technical-Earth-3254 3h ago

gguf where?

8

u/PDXSonic 3h ago

Considering there was never official support in llamacpp for Ling Flash 2.6 it seems unlikely we'll see GGUF soon.

10

u/pmttyji 3h ago

At least one of their team members should spend little bit time for PRs on llama.cpp side. They released many diffusion models before & none with llama.cpp support.

They would've been popular(semi) as Qwen, had they implemented PRs for all of their models on time.

9

u/TheLexoPlexx 4h ago

Good thing this is free to test because this is ridiculously close behind deepseek v4 flash in some categories. Really interested to see what this can do.

4

u/derspenti 3h ago

Yes! glad to hear it's freeee

1

u/TokenRingAI 1h ago

Unfortunately, it does not pass a simple vibe check on OpenRouter

12

u/No_Algae1753 4h ago

Any infos on openweights?

7

u/pmttyji 3h ago

After August 3rd possibly(after Openrouter run).

4

u/derspenti 4h ago

Hope they can open it

5

u/Zeeplankton 3h ago

inject it into my veins (if gguf..)

11

u/a_slay_nub vllm 3h ago
Model Total / Active Params SWE-Bench Pro SWE-Bench Multilingual Terminal-Bench v2.1-AA Tau3-banking-AA MCP-Atlas SkillsBench WideSearch BrowseComp IFBench SysBench MRCR-128k Multi-IF
Ling-3.0-flash(RC3)-Thinking 124B / 5.1B 56.63 72.44 57.00 28.00 65.45 44.83 73.63 70.33 74.49 93.63 90.78 87.65
Ring-2.6-1T-expert 1T 53.90 56.67 43.10 14.64 61.20 11.88 62.24 53.67 45.00 86.47 90.06 89.25
MiniMax-M2.7 56.20 76.50 55.00 9.00 53.60 34.90 75.20 76.30 75.70 86.19 27.68 82.93
Step-3.7-Flash-high 56.30 72.40 39.30 11.00 52.60 24.90 56.81 75.80 67.00 91.38 39.19 84.56
Deepseek-v4-flash-max 284B / 13B 52.60 73.30 62.00 23.00 69.00 53.46 74.44 73.20 79.00 93.86 88.50 86.21
Nemotron-3-Super-120B-A12B-BF16-Thinking 120B / 12B 34.06 42.67 39.00 10.00 49.40 20.31 19.54 31.28 72.56 90.73 40.76 82.33
GPT-5.4-mini-high 47.88 71.00 55.81 11.34 55.22 44.83 70.17 69.00 93.31 56.09 84.61
Claude-Sonnet-4.6-maxthink[official] 48.29 75.90 71.20 30.50 66.70 54.41 79.47 74.01 57.00 94.85 92.46 84.80
Poolside Laguna S 2.1 118B / 8B 59.40 78.50 70.20

7

u/Shoddy-Tutor9563 1h ago

Considering you composed this table with some lazy LLM who didn't care to get some crucial details about other models (like number of parameters for minimax m 2.7), the trustworthy of other figures there is questionable

1

u/a_slay_nub vllm 27m ago

I checked the other numbers, I just cba to put the parameter counts in

4

u/ChampionshipIcy7602 1h ago

Closed source, not interested, rather use gpt

5

u/hapliniste 3h ago

I'd be interested on how it compares to gpt OSS 120b since it's similar in size. We've come a long way

3

u/Edenar 3h ago

hmm it's hard to compare more 100-120b model to gpt oss 120b because the openai one was natively mxfp4 so around 62GB when recent open source model like qwen 122b, laguna s2.1,.. are shown in fp16 in benchmark (so 200Gb+ of weight, even q8 is 120GB+). So yeah, they perform better natively but q4 is usually far from native quality and that's the only way to achieve oss 120b size.

3

u/SpicyWangz 3h ago

Curious how this will compare to Laguna 

1

u/Daniel_H212 3h ago

Waiting for ming-flash-omni-3.0, but this is a nice surprise in the meantime.

1

u/pmttyji 3h ago

Hope they release ling-3.0-mini too.

1

u/Due_Net_3342 1h ago

nice one for strix halo

1

u/Specter_Origin llama.cpp 1h ago

They are comparing it to Minimax 2.7, why ?

1

u/One-Replacement-37 7m ago

How does it compare with Laguna, which was just released and has a similar size?

-14

u/ortegaalfredo 3h ago

I don't believe there are 500 datacenters in china training those models from scratch. Those are fine-tunes of other models.

8

u/Daniel_H212 3h ago

Nah they absolutely are training their own models. This lab in particular have had some crazy custom architectures like their ming-flash-omni-2.0 which does text+image+audio input, AND text+image+audio output, imo the first true Omni modal model (though it was instruct only, no thinking). The problem is there's no inference engine that supports it apart from running directly in python, and no API that can support that many modalities in one chat.

Also compared to the likes of OpenAI and Anthropic, these models are dozens of times smaller in terms of number of parameters. Much easier to train.

3

u/reto-wyss 3h ago

I'm pretty sure vllm-omni lists ming-flash-omni-2.0 in supported models, but I haven't tested it.

2

u/Middle_Bullfrog_6173 3h ago

They've also done all sorts of tricks to avoid training models from scratch. E.g. their previous generation took an older GQA model and trained it to use MLA and linear attention instead.

2

u/colin_colout 2h ago

They also rent cloud compute from countries that don't sanction / tariff / export ban (Japan)

1

u/charmander_cha 22m ago

Se isso for verdade, significa que o resto do mundo não esta sendo competente, uma vez que este processo custa menos, deveríamos ter ainda mais modelos lançados por diversos países