DraftReviewPublishedArchived

Free big models in the past 6 days: 5 have died and 3 have survived

Six days ago, I counted 18 free models that could still be paid prostitutes for free. Today, I ran it back to 15. But it's not as simple as three missing: five dead, three came alive, and one new, and the one that came alive is now the fastest in the game. This time, I added a repeat call to all entrances that returned to 429. I played 4 rounds every 45 seconds. As a result, Cloudflare and Groq both reversed their convictions. If I played only one round, the two fastest ones today would be sentenced to death by me. In addition, MiniMax has removed the entire basic line at the free level, including the 1049k context.

By Joker09/08/20265 min

Six days ago, I counted 18 free models that could still be paid for prostitutes. Ran it again today, 15.

But this is not simply three missing. In the past 6 days, 5 people died, 3 people came back to life, and 1 new one emerged. Among them, the one who came back to life is now the fastest in the audience.

!

first said something that almost made me write wrong

For the first round of scanning today, Cloudflare returned 429.

According to my past habit, I should write it down in the "hanging" column. It was still fine for 2108 milliseconds 6 days ago. I just used it to hit 15 times in a row yesterday, and I didn't miss once. Suddenly 429, it looked like I had exhausted the quota.

I called again after a few minutes, three times in a row, and everything was normal. Another round was repeated every 45 seconds, and a total of 4 rounds were repeated, 4 times were fully cleared, with the fastest speed of 1620 milliseconds.

That 429 is an instantaneous current-limiting window, not a hanging window.

After the same set of repeated matches, Groq also reversed the case. On September 2, I measured it Rate limit reached, and I put it in the hanging column at that time. Today's first call was still a current-limit. I played for 4 rounds, 4 all-passes, 579 milliseconds.

So this time I added a repeat call to all entrances that return to 429 or current restriction: one round every 45 seconds, four rounds. The result is this:

inletthe first roundRepeat 4 roundstrue state
Cloudflare4294/4 all-opennormal
Groqcurrent limiting4/4 all-opennormal
NVIDIA MiniMax-M34290/4I really can't use it
Cerebraspaywall0/4I really can't use it
Mistralsubscription level0/4I really can't use it

If only one round was played, Cloudflare and Groq would both be sentenced to death. And they are exactly the two fastest ones today.

The 429 on the free layer and the 429 on the paid layer are not the same. Paid-level current limiting usually means that you really exceed your quota. Paid-level current limiting is often just a moment when the machine is busy, and it will be fine after a while. The differentiation method is stupid but effective: play again every few dozen seconds and play more rounds.

5 of the ## were suspended, but the entire line of MiniMax was withdrawn

!

MiniMax-M3: It was still adjustable from the NVIDIA entrance 6 days ago, 17285 milliseconds, slow but working. Today lasted 429, playing 4 rounds 0 - 4.

minimax/minimax-m3:free: disappears directly from OpenRouter's model list. Let me add another word to this. Six days ago, it was the one with the largest context in the entire free list,1049k, just over a million yuan. If you want to put an entire book in and ask questions, it is the most suitable one. It's gone now.

minimax/minimax-m2.7:free: Also disappears from the list.

With two entrances and three models, the entire line of MiniMax is basically removed on the free floor.

poolside/laguna-xs-2.1:free: 1009 milliseconds 6 days ago, which was the fastest one in the :free at the time. Today is Provider returned error.

Google/gemma-4- 31b-it:free: It was available in 1666 milliseconds 6 days ago, but today it was also an upstream error. What's interesting about this is that another model of the gemma-4, 26b, has been reported incorrectly since August 28, when the 31b was good. Now they are both down.

By the way, the total number of models superscript :free on OpenRouter is also shrinking, from 18 six days ago to 16 now.

3 came alive

Groq oss-120b, 579 to 851 milliseconds, 4 rounds of all-pass. As mentioned above, I wrote it down in the hanging column six days ago, and this time it came back by itself.

Smart Spectrum GLM-4.7-Flash. Six days ago, it expired in 120 seconds and failed in 3 retries. Today, the result will be available in 19191 milliseconds. 19 seconds is indeed slow, but it is a qualitative change from "completely incomprehensible" to "can produce results".

poolside/laguna-s-2.1:free, upstream reported an error 6 days ago, today 2609 milliseconds. Note that it is of the same factory and the same series as the laguna-xs hanging above, one is alive and the other is dead. The usability of different models in the same house is completely independent. This rule has been established since the end of August.

There is also a newinclusionai/ling-3.0-flash-sante:free, 1319 milliseconds, which was not on the list 6 days ago.

also has a third state, which is more difficult than hanging

Free models do not only have two states: "usable" and "unusable". There is also one thatis nominally alive, but slow enough that you don't want to use it.

Nemotron-3-ultra-550b is a typical example this time.

inlet09-02todaymultiple
OpenRouter2629ms37407ms14 times slower
NVIDIA Direct Connect7007ms35328ms5 times slower

It reported no error, and the interface was not removed. If you adjust it, it will reply to you, but it will only take more than half a minute.

Two completely different entrances are slowing down at the same time, which means that the problem lies with the model itself, not the routing or current restriction of a particular home.

I set the threshold for availability at 35 seconds this time. If it exceeds 35 seconds, it will be unusable, so both entrances have been cleared out. This 35 seconds is my own standard, not an official standard. But if you really wait for it in code for 37 seconds, that link will basically be useless.

QKPFX4 The fastest position in QK has been changed four times a month

It's quite interesting to link up the records for this month:

  • August 17:Cerebras gpt-oss-120b, 593 milliseconds
  • August 25: Cerebras hit paywall, topof Groq, 545 milliseconds
  • September 2: Groq current limit,smart spectrum GLM-4-Flash top, 788 milliseconds
  • September 8:Groq is back, 579 milliseconds, back in first place

The first three times were "the previous one died, the next one was on the top." This time, the one who died came back and took back his position.

This rhythm shows one thing: the ranking of the free layer has no inertia. You remember the fastest this month, but there is a high probability that it will not be it next month. If you really want to use it, don't write a certain one in your code, write a list that can automatically switch.

!

QKPFX9 15 QK still available now

Five directly connected: Groq oss-120b (579ms), Smart Spectrum GLM-4-Flash (1040ms), Cloudflare oss-120b (1620ms), Kimi k2.6 (3910ms), Smart Spectrum GLM-4.7-Flash (19191ms).

OpenRouter superscript 10 :free, arranged according to speed: ling-3.0-flash-fin(1192ms)、ling-3.0-flash-sante(1319ms)、nemotron-3-super-120b-a12b(1519ms)、nemotron-3-nano-omni-30b-a3b-reasoning(1721ms)、lfm-2.5-2.6b(1884ms)、north-mini-code(2177ms)、nemotron-3.5-content-safety(2602ms)、 laguna-s-2.1(2609ms)、dots-3-note-preview(2787ms)、nemotron-3.5-lightning(16568ms)。

For long context, nemotron-3.5-lightning is nominally 1000k, and dots-3-note-preview is 512k. The former takes 16 seconds, but it can still be used after enduring it. It turns out that the 1049k minimax-m3 is gone.

How to match ##

don't write a dead one in the code. The first place changed four times this month, and none of them was right. Make a list, and if you fail, cut down.

429. Don't think it's dead. Try again first. Today, Cloudflare and Groq both reversed the verdict in this way. They gave up once a current limit, which is equivalent to throwing away the two fastest ones in vain. The cost of retreating and retrying several rounds is very low.

Count "too slow to use" as failure. It is not enough to just judge whether there is an error. The 37-second response like nemotron-3-ultra is completely normal at the interface level and has long been abandoned in business. Add a timeout to your call, and cut the next one if it exceeds it.

Different models from the same house should be tested separately. One of the laguna series lived and the other died, and the two gemma-4 models fell down one after another. This kind of thing has been seen several times in this month. Don't cross off the entire house just because one model doesn't work.

few sentences boundary

This is a snapshot of an account at a point in time. The amount of the free layer is calculated based on the account, and your result may be different from mine, especially in terms of current restriction.

I set the 35-second availability threshold. If you change the threshold, the available quantity will change.

I only tested whether it could return normally, did not measure the output quality, and did not run through the quota to see how much money I could pay for.

In the "Hang Up" column, NVIDIA Nemotron-49B is officially retired. The return states that the end of life will be August 26, 2026. This is certain; pay walls, subscription levels, and limit current limits. It may be good to change your account.

Each family's strategies have been changed very frequently. In the six days from September 2 to today, 18 have changed into 15, with 5 hanging and 3 active round-trips in between. The shelf life of this list is only about one week.

QUEST COMPLETEREWARD: +30 XP, +1 LEGENDARY ITEM
Build Progress100%
No signal
PULSE
0PULSES