Who can use the free big model directly? Who has to pay first?
After a few days, I reviewed the batch of entrances in my hand: 6 that could directly call up the results, 3 that had to be paid for first, and 1 model was missing. The fastest Groq is 605 milliseconds, and the slowest NVIDIA is nearly 32 seconds, which is 53 times worse, but they are all the same on any "still free sex" list. In addition, Kimi was more than three times slower, dropping from 3.9 seconds to 13.9 seconds. I retested it for 5 rounds to confirm that it was not shaking. There are also three people who are asking for money, but one is that the quota is used up, the other is that the card has to be tied at the beginning, and the other is that the subscription level is not enough, so the handling method is completely different.
I haven't tested it for a few days, and today I went through the entrances in my hand again.
Let's put the conclusions here for now: 6 models can be directly called up, 3 models can be used by paying for first, and 1 model is no longer available.
Another change I didn't expect was that Kimi was more than three times slower.
can directly use 6
| inlet | today | last time |
|---|---|---|
| Groq oss-120b | 605ms | 579ms |
| Cloudflare oss-120b | 1204ms | 1620ms |
| OpenRouter ling-3.0-flash-fin | 1248ms | 1192ms |
| Smart Spectrum GLM-4-Flash | 2297ms | 1040ms |
| Kimi k2.6 | 13921ms | 3910ms |
| NVIDIA NIM | 31962ms | 35328ms |
The fastest is 605 milliseconds, and the slowest is nearly 32 seconds, a difference of 53 times.
These six will be counted as the same item in any "list of free prostitutes" and are all called usable. But if you follow the list and pick one at random, picking Groq and picking NVIDIA are completely different experiences.
The first four are all within 2.3 seconds, so there is no problem to get on with proper business. Kimi has to wait more than ten seconds. NVIDIA's 32 seconds are useless if you put them into any process where someone is waiting for results.
Kimi is more than three times slower and stable
When the first round scanned for 22.5 seconds, I thought it was shaking, because it was still 3.9 seconds on September 10.
So 5 separate rounds of retest were repeated, each round separated by 15 seconds:
13882ms、12470ms、13921ms、16633ms、15418ms。
All 5 attempts were successful, with a median of 13921 milliseconds, and the fastest was 12.4 seconds and the slowest was 16.6 seconds.
It's not shaking, it's slowing down steadily. From 3.9 seconds to 13.9 seconds, more than three times.
When looking at long-term records, I found a pattern: these entrances tend to slow down first before stopping taking services. It's not that Kimi wants to stop, it's fine now, and it hasn't missed once in 5 calls. It's just that if you scheduled it in your process, the time limit should now be calculated again.
has to pay for three items first, three methods to pay for them
These three errors are all about money, but they mean completely different things.
Cerebras: I'll let you use it first and ask for money when you use it.
Payment required to access this resource. Visit your billing tab.
It runs out of free credit and hits a pay wall. This is best understood and the most common.
SambaNova: You have to tie the card from the beginning.
A payment method is required. Add one at ... to continue.
Note that this is not the same thing as Cerebras. It doesn't ask you to pay after you use it up. It's because you don't tie the payment method and can't adjust it once. For people who only want to prostitute for free, this family has not been on the list from the beginning.
Mistral: This model is not in your range.
This model is not available in your subscription tier
This is a hierarchy, not a quota. It doesn't matter how much you use, but the model itself requires a higher subscription level. Maybe a smaller model will work.
It is practical to distinguish between these three types: those who have used up their quota will be reset or changed homes; those who need to bind the card will be crossed out directly; and those who have insufficient subscription levels can try changing the model.
There are two things in ## that don't count, which are problems with my own account
This time there are two other things that don't make sense: volcanic arks and silicon-based flows.
Volcanoes return account arrears, and silicon-based flows return insufficient balance.
These two items have nothing to do with the platform's free policy. My two accounts are empty.
I took them out separately and said that this kind of error statement looks very similar to the three above, both talking about money, but the nature is completely opposite: the above three are doors set by the platform, and these two are not my own payment.
It's worth keeping an eye on when looking at any list of "what's still available for free". When others tested that "a certain family cannot be used", it may be that the family has really changed, or it may be that his account is in arrears. These two situations cannot be distinguished among the errors reported, so we can only rely on the person who tested them to explain them clearly.
still has one model missing
DeepSeek-V4-Flash on ModelScope returns:
Model id : deepseek-ai/DeepSeek-V4-Flash , has no provider supported
There is no supplier. This is not a charging issue, it is that this model is no longer available on this platform.
, so how to choose now
Choose from the top four: Groq, Cloudflare, OpenRouter, and Intelligence, all within 2.3 seconds. Groq is still the fastest.
Kimi re-timed the clock before making a decision. It works, but the response in more than ten seconds is not the same as it was three months ago.
NVIDIA is only suitable for offline tasks. 32 seconds, no one can wait.
When you see "asking for money", classify it first. The three processing methods of running out of quota, binding card, and subscription level are completely different.
There is another layer: is it the threshold of the platform, or is it the arrears of your own account. If you make a mistake, you will either wait in vain or change your home in vain.
A few sentences
A snapshot of one account and one point in time. The free quota is calculated by account number, and your number will be different from mine.
It only covers these entrances equipped in my own gateway. It is not a full-market census. There must be something I didn't measure.
Only the normal return and time consuming were measured, but the output quality was not measured.
Speed is affected by the network and platform load at that time, and single figures should not be regarded as long-term levels. I only dared to write Kimi's question after 5 rounds of retest, and the other few were all single times and are for reference only.
The two issues of volcano and silicon-based flow are my account's own problems and do not constitute any judgment on these two free policies.