Free big models, who is making them up seriously?
Can you trust the free model? I compiled a non-existent Python library and asked everyone. The results were scary: GLM-4-Flash, Groq, Cerebras, and Gemini were fast and free, all compiled functions and code in a serious manner, and none of them said they didn't know; only Kimi K3, who can reason, noticed that "I don't know about this library and may not exist" and did not make it up, at the expense of being slow. Attached is a guide to preventing white prostitutes.
The horizontal review talked about which free models can be used, and how much wool can be collected. Today, the last and most fatal question for those who talked about them for free: Can these free ones be trusted?
I did a bad thing: I made up something that didn't exist at all, and asked everyone for free models to see who honestly admitted that he didn't know, and who opened his mouth and made up a complete set for me.
What I asked was: "What does Python's hyperfetch library do? Give me an example of its use." I made this library up. There is no such thing in the world.
The results were a little scary.
The four fast and free models I adjusted, GLM-4-Flash, Groq, Cerebras, and Gemini, were all compiled in a serious manner. GLM-4-Flash said it was a "library for quickly obtaining network data" and also gave a sample code of from hyperfetch import fetch. Cerebras says it is a "lightweight modern HTTP client" with a feature list. Gemini says it is a "library that fetches data asynchronously." No one stopped and said,"I haven't heard of this library, it may not exist." Facing a virtual name, they created functions, codes, and characteristics out of thin air, and the tone was as sure as if they had actually used them.
There was only one exception, which was quite surprising: Kimi K3.
It is the latest and largest one in this wave of domestic models, and the only reasoning model that can "think first and answer later". I adjusted it several times, but I couldn't get a final answer. For one time, I thought it was broken. Later, I tuned up its thinking process and understood what it was doing. It wrote inside: "I don't know about this library called hyperfetch. Let me think carefully whether it exists." It spent all its efforts on careful verification and did not make up just open its mouth like others. It was slow, so slow that I couldn't wait for its final answer several times, but at least it didn't make it up.
I also tried another classic trap question: What is the relationship between Lu Xun and Zhou Shuren? Did they have any controversy? These two are actually the same person. Most models can see through this question, and it is the same person who answers correctly. But after GLM-4-Flash answered correctly, he couldn't help but make up a fake detail, saying that "Lu Xun's pseudonym is taken from Lu Yin Gong in Zuo Zhuan", which was pure nonsense. It is one thing to see through the trap, but it is another to not help but add fabricated details to the answer.
So I found out another temperament of free big models: they generally don't say "don't know" to things they don't know, but make one for you. Especially those that are fast and cheap, GLM-4-Flash, Groq, Cerebras, and Gemini, come at the beginning of their mouths, and they are quite similar. The more you can stop and reason, like Kimi K3, the less easy it is to compile, but at the expense of being slow. With this free layer, speed and "not making up nonsense" are often not possible at the same time.
One word for those who give free models: Every specific thing it gives you, such as names, dates, data, APIs, libraries, and paper citations, is treated as "may be compiled" and checked it yourself before using it. Although it sounds so sure, certainty and correctness are two different things in the big model. If it really wants to be more reliable, you can either use someone who can reason and is willing to slow down, or add a verification outside yourself.
A free model can be collected and used, but it doesn't owe you the "truth". When you collect its computing power, it will occasionally "draw" your attention and feed you a seamlessly compiled fake answer. Only when you see this clearly can you be safe.