How many times can a free big model prostitute for free? Up to 300, minimum 3
Before writing a free large model, I always explained at the end of the article that "I didn't run through the quota". This hole has been left for more than a month, and today I make up for it: I adjusted every entrance continuously and kept adjusting it to stop me. As a result, the gap was so big that I froze for a moment. I played Wisdom 300 times without saying a word. Kimi kicked me out the fourth time. I also found something in the middle that I didn't realize before: the 16 free models on OpenRouter shared the same daily quota, about 50 times. After playing, the 16 models were hung together, and the models that had not been touched were also blocked. Cloudflare's 10000 neurons are also account-level.
Before writing a free large model, I would honestly confess at the end of the article every time: I only tested whether it could be adjusted, and I didn't run through the quota.
This pit has been left for more than a month. Make up for it today.
The method is very simple and crude: adjust each entrance continuously until it stops me.
As a result, the gap was so big that I was stunned. They are also free. Some I called 300 times and it didn't say a word, and some kicked me out the fourth time.
Look at the numbers first
| inlet | How many times have I been stopped? | median delay |
|---|---|---|
| Smart Spectrum GLM-4-Flash | 300 times without being stopped | 1498ms |
| Cloudflare oss-120b | 165 | 1975ms |
| OpenRouter ling-3.0-flash-fin | 39 | 1081ms |
| OpenRouter nemotron-3-super | 17 | 1929ms |
| Kimi k2.6 | 4 | 4861ms |
Wisdom's 300 must be made clear: that is the upper limit set by me, not its upper limit. It didn't respond after running 300 times, so its real limit was not measured this time. I asked 300 questions in a row in 550 seconds, but didn't stop me once.
QKPFX1 One of the most valuable items in QK today: a bunch of free models, sharing a quota
The two rows of numbers in OpenRouter are broken 39 times and the other breaks 17 times.
My initial understanding was "this model has less quota, and that model has more quota." Until you see clearly the error:
Rate limit exceeded: free-models-per-day
It is free-models-per-day, not a limitation of a certain model.
38 plus 16, exactly 54 times.
Then it needs to be verified. I changed to three models that I had not touched today and typed again:
| model | results |
|---|---|
| cohere/north-mini-code:free | Same free-models-per-day |
| liquid/lfm-2.5-2.6b:free | exactly the same |
| nvidia/nemotron-3.5-content-safety:free | exactly the same |
Sit down. The 16 models labeled :free on OpenRouter share the same daily quota, about 50 times.
What this means is: You open OpenRouter and see more than a dozen free models, and you think the resources are so rich. They are actually followed by the same counter. The number of times you spend on Model A, Model B is gone. After playing 50 times,all 16 will be hung together, not just the two you have used.
Cloudflare is the same routine. When it stopped me, it said:
you have used up your daily free allocation of 10,000 neurons
I changed it to two other small models and typed it again, and the return was exactly word for word. 10000 neurons are account-level, not one for each model.
The 10000 neurons were tested 164 times, with an average of about 61 each time.
So here's an anti-common sense conclusion: the number of models does not equal the number of available models.
There is only one free model of Zhipu, and I called it 300 times without stopping. There are 16 OpenRouters, which add up to 50 times a day. One unlimited one is much more practical than sixteen shared 50 times.
Groq If I want to talk about this alone, I almost got it wrong again
Groq's score was 0 success times in the test, and he was blocked on the first time.
But it was still fully cleared in 4 rounds and 579 milliseconds yesterday. If you write the number 0 success directly, it will tell you that Groq cannot be used, which is wrong.
So repeat the game, one round every 40 seconds, five rounds:
| rounds | results |
|---|---|
| round 1 | Success, 619ms |
| round 2 | Current limiting, 226ms |
| round 3 | Success, 642ms |
| the 4th batch of | Current limiting, 225ms |
| round 5 | Success, 655ms |
Those who succeeded were all in 620 to 655 milliseconds, and those who were blocked were all in 225 milliseconds per second. Stop it every other time.
Groq is not a "quota has been exhausted", buta continuous fine-grained current restriction. It is always available, but the actual throughput is discounted at half. In the burst test, it succeeded 0 times because it had just played a round before and fell in the current-limiting window.
This is the third time in the past few days. The day before yesterday, I almost ruled Cloudflare failed in its first round of 429. Yesterday, I checked the records and found that I had misjudged Groq twice. If you report errors when reporting current limits on the free layer, you will basically be wrong if you draw conclusions without calling them.
Kimi 3 times per minute, the results are the same for both measurements
The fourth time was 429, and the return note stated that it was 3 times per minute.
I tested it once on August 27, and it was the fourth one to be stopped. Half a month later, the results were exactly the same, indicating that this was a stable strategy, not just in time that day.
two restrictions, the response is completely different
This time, there are two types of methods to block people. When you mix them together, it will be chaotic.
Rate type: The maximum number of times in a period of time. Kimi 3 times per minute, Groq limits current every other time. The characteristic is that thetotal amount is not capped, and it can be used up at a slower pace. Kimi, if you ask every 20 seconds, you can ask more than 4000 times a day.
Total amount type: This is the amount per day until it is used up. Cloudflare 10000 neurons, OpenRouter 50 times. The characteristic is that itis useless to slow down, and you can only wait until the next day to reset after playing.
The cost of confusion is quite practical: you think you can save it by slowing down, but it turns out that the company is a total amount type and uses it slowly so many times; or you think you give up after the quota is exhausted, but it turns out that the company is a rate-rate type, wait 30 seconds and you can continue using it.
, so how to match it
To run batch tasks, use wisdom. The only one that didn't measure the upper limit this time was that 300 consecutive rounds did not have any interception.
Don't use OpenRouter as a quota pool. Its value is "You want to try a specific model", not "There are a lot of free credits here." 16 models were shared 50 times, and each model was equally shared 3 times.
Cloudflare is suitable for daily scattered use. 164 times a day, it is enough for one person to ask and answer questions normally, but if you use it to run batches, it will be exhausted once.
Kimi can only serial, leaving 20 seconds between two requests.
Groq's current limit should not be suspended. There will be retries. The successful times were all in just over 600 milliseconds, which is the fastest in this batch.
few sentences boundary
One account, one day's test. The free quota is calculated by account number, and your number will be different from mine.
The 300 times of Smart Spectrum is the upper limit set by me, not its upper limit. We couldn't measure how much it gave this time, so we would find an opportunity to increase it later.
Max_tokens is uniformly set to 150, so the number of words output is subject to this limit and cannot be used as a basis for "how many words can be written".
The number of neurons Cloudflare consumes each time is related to the model and output length. The number of 61 is the average value under this configuration this time, and it will change with a different model.
Finally, to be practical: this kind of test will use up up the day's quota and affect other tests that day. This is also why it took more than a month to do it.