DraftReviewPublishedArchived

The unit price of Mixed Hy4 is 6 times more expensive, but the measured bill is only 1.8 times more expensive.

Tencent Hunyuan today released and open-source a new generation of Hy4 preview. The parameters of the press conference can be seen elsewhere. What I care about is how expensive and stronger it is than the previous generation. The input unit price of Hy4 is 6.3 times that of Hy3, and the context gives 1.04 million tokens. The same question is run on both sides: Hy4's three questions total 37 seconds and US$0.0044, and Hy3's 50 seconds and US$0.0024. The unit price difference is 6.3 times, and the actual bill is only 1.86 times, because Hy4 thinks much less. The previous generation burned all the 2000 budget on the issue of writing code. 4556 words were written in the thinking area, and not a word of the text was output.

By Joker08/28/20265 min

Tencent Hunyuan today released and open-source a new generation of Hy4 preview.

You can see what the press conference said, how many parameters are, and how high the score is. What concerns me is another question: how much more expensive and stronger it is than the previous generation.

Because both numbers can now be found. The Hy4 preview is already on the shelves, and the previous generation Hy3 is still on the shelves. The price is open. I just need to run the same question on both sides.

Let's talk about the price first.

Hy4 inputs $0.834 per million tokens, and Hy3 is 0.132. 6.3 times more expensive. The output side is 2.501 to 0.528, which is 4.7 times more expensive. The context gave 1.04 million tokens, which is four times that of Hy3.

Then the question is very specific: Is it worth it six times more?

!

stepped in the pit first

The first call, I gave max_tokens: 120 as usual, and the bodies returned by both models were empty.

Only by looking at usage can you understand:

completion_tokens: 120
reasoning_tokens: 120

All 120 tokens were eaten up by the thinking process, and not a single word of the text began to be generated.

Hy4 and Hy3 are both inference models. If you adjust it according to the pose of a normal model, give it a small max_tokens and a read-only content field, you will get an empty return, and you will think it is dead.

I stepped on this pit once on other models, but I didn't expect to hit it again on the day of release. So to adjust these two, two things must be done: give enough max_tokens and read the reasoning field at the same time.

After changing it to 2000, testing will be officially launched.

three questions

The first question is a trap reasoning question: the water lily doubles every day and fills the pond on the 48th day. When will it grow half?

Both were correct, day 47.

But the process is much worse. Hy4 used 118 thought tokens, 4.5 seconds;Hy3 used 484 thought tokens, 8.6 seconds. Also correct, Hy4 only had a quarter of the amount of thought that the other person had.

The second question is a Chinese copy. Write a main picture copy for the noise-canceling headphones. Within 30 words, it is clearly required that empty words such as "opening a new era" are not allowed.

Hy4:"Put on the subway and be instantly quiet. You can listen to it for a week after charging it once!" Hy3:"It takes a month to recharge the electricity once. Wear it on the subway and you can sleep quietly."

There was no empty talk, and both passed this point. Hy3's sentence,"It's so quiet that you can fall asleep", has a better picture, but "it takes a month for a charge" is a bit overblown. If you press this release, something will happen. Hy4 says "listen for a week" more practical. I decided that this question was tied with Hy4.

Write the code for the third question, an IPv4 verification function.

Hy4 took 252 thought tokens and 8 seconds to deliver the code completely.

Hy3 is overturned.

!

Its thinking process wrote 4556 characters, consuming all the 2000 max_tokens, and not outputting a single word in the main text.

Keep the money, but don't do the work. This is what I want you to see most: When the reasoning model gets out of control, it doesn't give you a wrong answer, it gives you a blank, and the bill still arises.

calculates the account

After completing the three questions, the total is like this:

Hy4:37 seconds,$0.004418. Hy3:50 seconds,$0.002378.

The one with a unit price of 6.3 times more is only1.86 times moreexpensive. Moreover, he was one-quarter faster, and he even solved one more question.

!

The reason is the previous thinking tokens: trap questions 118 versus 484, code questions 252 versus 2000. Hy4 thinks much less about each question than Hy3, which is enough to erase most of the six-fold difference in unit price.

This matter puts forward a general judgment, which I think is more worth remembering than the Hy4 itself:

For the inference model, horizontal price comparisons based on unit prices are ineffective.

You see that model A costs 1 yuan per million tokens and model B costs 6 yuan. The intuition is that A is six times cheaper. But if B only requires one-quarter of the amount of thought to solve the same problem, the real gap becomes less than twice. And if A still burns out the budget on complex questions like this time's Hy3 and then hands in vain, the money will be a pure loss, because you will have to try again.

So when selecting a model, what you should compare is not the unit price, butthe same question in your real business. Run it on both sides to see the final bill and success rate. This test took less than ten minutes to do.

few sentences boundary

Hy4 is a preview version, and the performance and pricing of the official version may change. This must be made clear.

I only ran three questions, once for each question. The sample size was very small, which could only explain the tendency and could not be regarded as the evaluation conclusion. In particular, the failure of Hy3's code problem is an observation this time. It does not mean that it cannot do code. Changing the prompt words or giving a larger budget may be excessive.

The price is taken from the quote returned by the OpenRouter interface, which is the price of this aggregation platform, which is not the same as the pricing of Tencent's official direct connection.

I didn't verify the 1.04 million tokens in the context, just the nominal value returned by the interface. Bidding for 1.04 million yuan and actually eating 1.04 million yuan are two things. I didn't do this step.

Finally, this article is not recommending or criticizing anyone. On the day of release, everyone can see the same bulletin. I just ran through the parts that I could run by myself and showed you the accounts.

QUEST COMPLETEREWARD: +30 XP, +1 LEGENDARY ITEM
Build Progress100%
No signal
PULSE
0PULSES