Two more people have been hanged, and now these 18 are the only ones who can still prostitute for free
Last month, I said Groq was the fastest free entry, 545 milliseconds, and suggested ranking first. If we run again today, it will limit current. Cerebras, which previously sat in this position, hit the paywall at the end of August, and Mistral Large also died this time. In three weeks, the fastest position has been changed three times, and the top two have been dropped. Now the first one is Smart Spectrum GLM-4-Flash, 788 milliseconds. The whole number of disks ended like this: the number of directly connected households dropped from 8 to 6, while the number of free pools in OpenRouter increased from 10 to 13. All 18 entrances that can be used now are in the article. It is recommended to take screenshots directly.
Last month, I said Groq was the fastest free large model entry at 545 milliseconds, and recommended that everyone rank it first.
If we run again today, it will limit current.
Cerebras, which had held this position before it, hit the paywall at the end of August. Going forward, Mistral Large also hung up this time, returning "Subscription level not supported".
In three weeks, the fastest position was changed three times, and the top two places were dropped.
Now sitting first is the Intelligent Spectrum GLM-4-Flash, 788 milliseconds. It was still 1590 milliseconds a month ago, and once slowed to 2546 in the middle, but this time it accelerated to 788.
After a round of trouble, the last one was the domestic, domestic directly connected, and permanent free one.
I will give you all the ones that can be used now. There are 18 in total. It is recommended to save the screenshots directly.
directly connected to each family: 6 left
This is the entrance that you can directly adjust with each key.
| inlet | time-consuming | 8-27:00 |
|---|---|---|
| Smart Spectrum GLM-4-Flash | 788ms | 2546ms, three times faster |
| Cloudflare oss-120b | 2108ms | 4205ms, twice as fast |
| OpenRouter Nemotron | 2629ms | 1437ms |
| NVIDIA Ultra-550B | 7007ms | 1377ms, 5 times slower |
| Kimi k2.6 | 8970ms | 11716ms |
| NVIDIA MiniMax-M3 | 17285ms | The current was limited before, but now it is restored |
Within a week, this gear dropped from 8 to 6. The three dead ones are:
Groq oss-120b, current limiting. It was the fastest in the game last week.
Mistral Large, replied,"This model is not available in your subscription tier." Pay attention to the nature of this sentence. It is not a current restriction or an arrears. It is that this model is not in your subscription level and cannot wait.
NVIDIA Nemotron-49B, officially EOL at 09:00 on August 26, and now returns directly to 410 Gone.
OpenRouter's free pool: 13, but it has become more
This file is a model with the :free suffix on OpenRouter, and all key can be adjusted.
| time-consuming | context | model |
|---|---|---|
| 1009ms | 262k | poolside/laguna-xs-2.1:free |
| 1135ms | 66k | liquid/lfm-2.5-2.6b:free |
| 1247ms | 256k | cohere/north-mini-code:free |
| 1332ms | 262k | nvidia/nemotron-3-super-120b-a12b:free |
| 1344ms | 1049k | minimax/minimax-m3:free |
| 1520ms | 128k | nvidia/nemotron-3.5-content-safety:free |
| 1666ms | 262k | google/gemma-4-31b-it:free |
| 1892ms | 262k | inclusionai/ling-3.0-flash-fin:free |
| 2404ms | 512k | dots-studio/dots-3-note-preview:free |
| 2555ms | 1000k | nvidia/nemotron-3-ultra-550b-a55b:free |
| 2597ms | 1000k | nvidia/nemotron-3.5-lightning:free |
| 2819ms | 256k | nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free |
| 3826ms | 197k | minimax/minimax-m2.7:free |
When I tested this batch four days ago, only 10 of the 18 were available. Today it became 13.
The three that came back: liquid/lfm-2.5-2.6b and google/gemma-4-31b-it were both Provider returned errors last time. nvidia/nemotron-3.5-lightning timed out of more than 35 seconds last time. This time, they came back in 2597 milliseconds.
There are still five that don't work. The two ones for thinkmachines are marked free but are only open in specific agent frameworks. The standard interface cannot be adjusted, so don't waste time.
By the way, remember one detail: Google's two gemma-4s, 31b came alive this time, and 26b was still a Provider error. The usability of the same family, the same series, and two adjacent models is completely independent. Judging whether it can be used can only be tested one by one, and cannot be inferred based on the entire manufacturer.
is the reverse phenomenon
Putting the two gears together, the changes in the past few days are:
Directly connected to each family: 8 → 6, a net decrease of 2. OpenRouter free pool: 10 → 13, a net increase of 3.
One side is shrinking and the other is expanding.
It makes sense to think about it. The free layer of a single platform is pure cost, and it is closed at any time. Cerebras 'pay wall, Mistral's subscription level, and Groq's current restriction all happened in the past few days, and each party has the final say.
The free pool of the aggregation platform is provided by each upstream family in turn. If one family fails, the others are still there, and the overall situation is more stable.
So I adjusted the order of configuration this time: the main force hangs on the aggregation layer, and the direct connection only leaves the most stable one as the bottom.
Specific to today's data, I will arrange it like this:
Use OpenRouter's free pool for daily calls. Select from the above 13. There are six or seven options for the period from 1009 to 1666 milliseconds, and hang three for rotation.
GLM-4-Flash. 788 milliseconds, domestic direct connections do not bother the Internet, and are forever free, and it is the only one that has not had an accident in this month.
For long contexts, there are three OpenRouter that give more than 1 million tokens. Minimax-m3, nemotron-3-ultra, and nemotron-3.5-lightning are all free.
Stop ranking top of Groq and Mistral. These two were just lost this week, and they will be put in reserve after they are restored.
few lines
This is a snapshot of an account and a point in time, September 2, 2026. The ones that are currently limited may be ready tomorrow, and the ones that can be used today may also be limited tomorrow. The shelf life of such lists is measured in days.
There are two types to distinguish between "dead" files: GitHub and Nemotron-49B are retired from the service itself and no one can use them; while arrears, subscription levels, and pay walls are strictly "I can't use this account", and you can be resurrected by changing your account or charging your money.
The context length is the nominal value returned by the interface. I did not pressure test one by one to verify the true available length. In addition, I only tested whether I could return normally, not the quality, nor did I run through the quota.