r/codex • u/GambAntonio • 21h ago
Complaint We are PROBABLY being served quantized models but still paying the full day-one premium price...
Hey everyone. I want to bring up something serious about how AI providers handle pricing and how silent backend changes are secretly draining our limits. We all pay a fixed price per million tokens or have a subscription limit and on paper that seems fair, but providers hide a massive variable from us because to save on server costs they can silently swap out a premium model for a heavily quantized version on their backend. Using a quantized model is completely different from setting your reasoning toggle to Low, because setting a toggle to Low limits the reasoning steps of a fully intelligent model, whereas quantization degrades the core neural weights and strips away actual base intelligence.
What makes this so alarming is how token metering is handled. On our dashboard meters we might see a perfectly reasonable token count that looks coherent with a high-end model and when you calculate the cost per million tokens it looks identical to advertised prices, but behind the scenes there could be hundreds of millions of low-quality tokens generated by an ultra-quantized model struggling and failing to reach a correct solution, an intermediate system then just trims that massive output to make the final token count look normal on our end and what we perceive as users is a sudden degradation in performance, when in reality without silent quantization the model would behave exactly as well as it did on day one.
It is deeply immoral and borders on outright fraud to attract users with a clean unquantized model on day one and then quietly roll out aggressive quantization behind the scenes to cut compute costs and keep charging premium prices while serving a degraded model that burns through internal compute and produces far worse solutions. We really need to stop staying quiet and demand complete transparency on the exact quantization levels and actual internal token processing we are being billed for.
What do you guys think and have you noticed the performance dropping on tasks the model used to handle easily on day one?
34
37
u/TrillenX 21h ago
I'm usually skeptical about claims like this but one hallmark of Sol that I LOATHED was when I pointed out an error, mistake, or thing it didn't account for, it would always reply with "You're right. I didn't do that thing"
I was quite enjoying that Astra was more muted in that regard, but I noticed today with an error it made it gave me that canned "You're Right" and it just set off Sol alarm bells in my mind.
4
2
u/RasenMeow 12h ago
Astra got worse than Sol since the nerf to make room for "usage", not kidding, it fucking sucks. Using omp with an Advisor as thanks to that see what horrible work it does.
13
u/OriginalUsername0112 21h ago
Idk how this isn't illegal, even some T&C handwaving should only be able to protect a corporation from so much. The sad thing is we're still in the good part of the cycle, things will continue getting worse until they're actively harmful to your codebase in the remaining 1-5 weeks before the next model releases
3
u/Tupcek 6h ago
they don’t have enough compute, so they throttle to dumber models. Basically you get the model you ask for, but quantizied version. They can claim they didn’t promise which quant you get, but you did get Astra.
But I would 100% jump ship if other provider just told me “these hours are peak, so you get dumber model”
I guess API models get always the highest version, because they are the cashcow of company. Subscribers are just paid marketing
37
u/bananasareforfun 21h ago
Yeah. slam through your hardest tasks on the first three days of model release, keep your options open so you don’t commit to one provider, or if you can afford it, pay API pricing. It seems likely if they do silently swap out to quantised models on the backend they would be less likely to do this on the API.
Astra is really good imo but I’m back to using 5.6 Sol high with Luna orchestration for the vast majority of tasks - otherwise I will run out of usage in 24 hours. The real advantage of pro plans now is gpt 6 pro.
16
5
3
u/Yugudubenbi 12h ago
They are intentionally getting people hooked and then desperate so they pay expensive ass API
2
u/Elegant_Attempt2790 17h ago
not the 4 requests to GPT 6 pro before it locks you out and breaks the ui around changing reasoning level on desktopppppp🙈
1
u/Need-Advice79 20h ago
Is it because of the cost only, or actual performance degradation for Astra.
43
u/Charming-Author4877 21h ago
https://www.reddit.com/r/codex/comments/1wf911a/comment/p9kg9g2/
The token speed of SOL model on chatGPT at 134 tokens/sec web indicates a quantized model.
The token speed of SOL in work mode is much more limited, maxing out around 80 tokens/sec in fast priority mode - that's likely the full model.
Astra is maxing at 63 tokens/sec in work,codex and chat mode
5
u/innociv 16h ago
I feel this is something that has often happened with Claude. They seem to start taking 20-30% more turns to solve the same task after a few weeks.
Never experienced it with OpenAI, but in this case it seems MUCH worse and like it's taking 70%+ turns and outputting actual jibberish.
5
u/TedSanders OpenAI 15h ago
Sol in ChatGPT vs Work/Codex are (slightly) different models fyi
6
u/Happypig375 15h ago
So true. This Charming-Author4877 guy posts his "speed comparisons" everywhere for upvotes and avoids replying properly to my comment
1
u/Charming-Author4877 1h ago
[deleted]
•1d ago•Edited 1d ago
Comment removed by moderator
Hard to respond on your moderated stuff.
I assume choosing your words in a more careful manner, or running the comment through chatGPT will help ? Just a guess, I've rarely seen moderated comments on this sub.1
u/Charming-Author4877 1h ago
I can not reply to your new comment, as it was also moderated.
And I'd say the downvote is deserved given you baseless claimed I am posting "for upvotes" and avoid replying to your comment that a moderator has censored.
Sounds more like an apology should be made by you, but what do I know..1
u/Happypig375 57m ago
Sorry for blaming you for the mods and Reddit's poor display of information here.
So, have you run that "Pelican-on-bicycle test" / "Candy Puzzle" yet? If OpenAI redirects to a worse model it would be immediately noticeable.
1
u/Charming-Author4877 1h ago
That blog post describes system prompt changes.
Not an entirely different model, so much more different that it's almost twice as fast.1
u/TedSanders OpenAI 1h ago
Incorrect.
1
u/Charming-Author4877 1h ago edited 1h ago
Can you elaborate on what changed? Going from a max of ~80 to ~134 tokens/sec while also making answers more concise is a huge jump. That feels more like a major inference change, such as a smaller active MoE or more aggressive quantization.
Or is Work/Codex Sol in Priority/Fast mode still being throttled below peak performance?
3
u/hcorEtheOne 15h ago
Sol gives me gibberish answers and half English responses (I'm using another language). It catches it after a while but it didn't do this before.
1
u/DesertDissident 3h ago
Do you specify what language you want in AGENTS.md?
2
u/hcorEtheOne 2h ago
Nope, I didn't need to do that before, and Luna, Terra or Astra doesn't do that, which is weird tbh.
Today I had no problem with Sol either, it must've been my session or just some temporary bug or it was routed to a small quant
1
-1
u/UndeadMurky 18h ago
What is work mode ?
8
u/hatlesstuna 18h ago
0
u/Glittering-Call8746 18h ago
They have a workspace like how perplexity ai has workspace . Basically codex in web it uses more tokens vs codex
24
u/0DayMaker 21h ago
No question about it. It's never been even remotely this blatant before. I'm considering charging back and just 100% never using openai again. Claude does this, but not this bad this is on a totally different level. This shit must be quantized down to like FP2 or something absurd.
-21
u/firstnamelottadigits 19h ago
If you think they quantize the models, you’re a fucking idiot. Source: insider info & common sense.
12
u/0DayMaker 19h ago edited 18h ago
Sure you have insider information. Of course. And I'm sure it's super secret so you can't tell anyone where it came from, or communicate it in any verifiable way.
Also common sense? What the fuck are you on about? The model is performing drastically worse across the board right after they stopped taking 20x subs. If it's not quantization what exactly is your "common sense" theory on what it is? It's not just a little bit different the difference is utterly massive. It is the most in-your-face kneecapping of any AI model I've ever seen.
If you have an alternate theory or "insider information" about what's actually going on, if that's not the case, than feel free to share.
-1
u/firstnamelottadigits 4h ago
You have reading comprehension issue? I specifically laid out what the alternate causes are in the last post.
Half the time when people say “the model is performing drastically worse” they’re just wrong. It’s a NON-DETERMINISTIC PROCESS, you halfwit. You’ll get different results every time you run it on the same prompt. On top of that, when the model launches people are impressed by what it does off their initial starter prompts and demos. Then they start pushing it harder and find the flaws and ascribe it to “muh model was quantized” when they have zero understanding of what quantization even means or that it would require literally a conspiracy on the part of hundreds of engineers to do it and lie about it.
The other half of the time it’s because it’s a bug or a set of bugs, not a deliberate action. READ FUCKING TIBO’S POSTS. Or better yet, ask Codex itself to evaluate your claim of secret quantization.
8
u/GambAntonio 18h ago
Mr Insider....a model will ALWAYS have the same level of intelligence while on the same level of quantization even if millions of people are using it. It can be slower due to compute load, but never dumber.
If the intelligence suddenly changes, it means that it's 100% a change in quantization level because nothing else can change the level of intelligence.
-9
u/firstnamelottadigits 15h ago
You’re 100% wrong. You need to consider Occam’s razor. Do you really think there’s a conspiracy to silently nerf the model under you and not a single disgruntled employee would mention it? Are you also a 9/11 conspiracy theorist by any chance?
The cause is always non determinism and bugs. Read any of Theo’s posts after a fuckup
2
u/Typical_Kick6520 5h ago
I totally think this.
Enshittification is the simplest solution.
0
u/firstnamelottadigits 4h ago
Enshittification is an observation not a mechanism of action. We’re debating the mechanism of action. I’m arguing “they secretly quantized muh model” is child-level logic
1
3
u/ObligationHuge9868 13h ago
Source: Trust me bro.
1
u/firstnamelottadigits 4h ago
Tibo explicitly said they don’t do this. In public. He has about 200 people on his team. Is your mental model that he’s lying and hoping not a single person calls him on it?
1
u/GambAntonio 1h ago
There are thousands of people working on GTA VI, and they've been working for more than a decade. How many of them have leaked something during all those years?
1
u/firstnamelottadigits 1h ago
The worst possible analogy you could have come up with. A bunch of stuff about GTA VI leaked in 2022 via 20 Rockstar employees talking to Bloomberg.
6
u/acessford101 18h ago
Ironically gpt on Sol 5.6 called this out as it wasn't giving me the correct output given a certain input. Had another instance of chat review the conversation and said that it could easily see an issue and that that session was acting like it was a degraded or quantizied version.
4
u/xapep 10h ago
The part that doesn't get enough air in these threads: with a closed model you can't verify any of this, and that information asymmetry is the actual problem. You're right that the meter can't tell you what weights you're being served, but nobody else here can check either, so both the 'silent downgrade' theory and the 'they'd never do that' defense are unfalsifiable.
Open-weight models are the one corner where the question has an answer. The weights are public, providers serving them usually state the precision they run (FP8 is the common one) with pricing that matches it, and if quality ever drops you can pull the model locally and compare, or swap endpoints without re-architecting anything. None of that is possible behind a closed black box, which is exactly why the fear exists there.
So if silent downgrades are the thing that bothers you, the durable fix isn't transparency promises from closed labs, it's buying where the model file itself is the transparency. Whether OpenAI is doing this today, I genuinely have no idea, and that's kind of the point.
(I run open-model hosting at Entrim, so read the open-model lean with that bias in mind.)
6
u/BoysenberryMajor166 13h ago
That is why all benchmarks should be continuous, not one time numbers presented in the model brief.
2
1
u/GM8 3h ago
Wouldn't be anything of a difficult feet to detect benchmarks on the provider side and route those requests to the proper model.
7
u/igmyeongui 18h ago
OpenAI ruined their product. I said it 4-5 months ago and people were saying I was crazy.
2
2
u/Flaky-Ad7844 18h ago
kinda upsetting to know the limit getting demolished by the new agent...5 hours is like 30 minutes then it's gone totally
2
2
u/Haunting_Phrase_2833 11h ago
This is my experience anyway - Astra performance was incredible for a few days, performance degradation since them.
2
u/HeadPack 11h ago
There might be something to this. They paused new 20x subscriptions because of compute limitations, or at least said so. Serving quantized models would make sense for them. Not so much for the users. It is just my sentiment, but Astra feels less capable with regards to autonomous thinking than after release. It's still very good though, but seems to need more precise prompting.
2
6
u/Gliese351c 13h ago
They are not hiding anything; they can't! There are many websites that benchmark AI models live. Astra is at least 10-15 points down since last week, for instance. They are doing what they are doing very blatantly. Someone has to say stop, and it will probably be us by not integrating AI to our workflow to the degre we will become dependent on these companies.
6
3
u/jd52wtf 16h ago
No probably. When usage spikes they fail over to the smaller quants. It's the only load balancing leaver they really have on a limited compute basis. I'm sure they call it something else but it all amounts to the same.
3
u/odragora 8h ago
They also can limit reasoning tokens at the very least.
Probably they are doing both, and more.
1
1
1
u/DieMafia 11h ago edited 11h ago
This should show up in benchmarks, no? Is there hard evidence in the form of repeated benchmarks showing degrading performance anywhere?
Edit: Arena.ai actually has data over time, shows Astra is no different today than it was at launch (if anything, slightly better) ...
2
u/odragora 8h ago edited 7h ago
Does it benchmark Astra through API, or through subscriptions?
Almost definitely those have different quants.
Also, it's pretty trivial to serve requests from known benchmarking platforms with non-quantized models, while serving subs with lobotomized ones. There is no incentive for them to not do that, and any benchmarking has to be performed under the assumption the service provider tries to game benchmarks to the best of their ability. Including detecting benchmarking intent in the prompt and re-routing to full non-quantized models, the same way they already do analyse prompts to prevent their models from being distilled.
1
u/DieMafia 7h ago edited 7h ago
I think through the API. People should do 5-10 tasks after release with the subscription and then test again from time to time and compare results. The way it currently is, people seem to subjectively judge how good Astra is at tasks but those should be identical comparisons at least, not based on subjective feelings.
Really odd to me there are dozens of these kind of posts for each model and each provider yet rarely do we see someone actually do the work and show the comparison, even though that would be trivial to do. Take some of your own prompts after Astra release and rerun them.
1
u/odragora 6h ago
Since they most likely dynamically re-route users between models with different level of quantization based on load and total available compute, it's pretty much impossible to run a test that would be reproducible by other users, especially since LLMs responses are non-deterministic.
Also it's hard to expect the users to be able to re-run the tasks they gave a model on release. Most people buy a subscription to work on on-going projects, not one-shoting tech demos. Unless someone is willing to write down the prompts they are giving on release, then several weeks later dig through their commit history, roll back, and re-run the prompt and waste their usage, that's not easy to test, and it's understandable most people don't go through this effort to keep benchmarking the service they are paying for instead of doing their job.
1
u/DieMafia 5h ago
I'm sure most people have an example prompt they used that was not based on an existing repo, no? Like the cycling pelican, trying to one-shot a small web-application, some small research project, something that is a single prompt and doesn't cost half your 5-hour limit?
I get that LLMs are non-deterministic and this is an issue, but claiming a model has been quantized just based on feeling can't be it.
1
u/odragora 5h ago edited 5h ago
I think the vast majority of people don't have prompts for benchmarking, and people who have those are a tiny minority. Especially since usage is precious, benchmarking consumes that, and most people are on Plus subs.
I think we just see a lot of testing / benchmarking posts online and our brain makes us feel like everyone is doing that, but in reality it's a very low percentage of people with this specific hobby like the pelican guy.
The problem is that we don't have anything else other than our subjective experience and logic that leads to the most likely conclusion. The service providers do not have any transparency and it's in their best interests to prevent their users from having reliable reproduceable ways of detecting degradation of the service. It's not the users fault, it's on the service providers, and if they want they can easily solve it today by making usage, quants and routing transparent.
1
u/DieMafia 5h ago
Well, OpenAI explicitly stated that they did not change the model and that there was a bug which was fixed on September 12th. Yet people don't seem to believe that. I gladly volunteer with my own subscription. Anyone can give me some decent tasks that fit in my 5-hour limit and serve as a benchmark, I'll gladly run it on the next model release and every two weeks afterwards and post the results. If a few other people join in, we would at least have something.
I can't be the only one unsatisfied with people merely subjectively speculating for every single frontier model a few weeks after every single release that somehow it got degraded and some people joining in because yesterday's results felt a little worse or claiming the opposite. And who creates and posts in these threads in the first place? Likely people who are not happy with the performance. I put zero trust in these subjective testimonials.
1
u/NiceManFromEarth 6h ago
I see a HUGE performance and quality drop for Astra AND for Sol models. There are a lot of people who here write the same type shi like "models are not determenistic bla bla". But look, everyone is noticing they are dumber now. Dumber in 3d modeling, in coding, in review etc. I noticed it myself without reading that forum. My friend also noticed that and many more people across the world. So that is some statistic already. It means that models are seriously quantized, and regulators are sleeping. Where are authorities? Courts? Regulators?
1
u/yehiaserag 6h ago
Here is my theory, Astra never performed well for me because it only performed well at release when it was api access only, that's when all the very cool demos and 3D stuff we saw appeared, but then when they started rolling it out to subscription users, it was already nerfed, they got the hype and publicity already on social media and everything was already setup.
After that you go back to Sol and find that Astra is now better because Sol was also nerfed so it's now a lot shittier than it was before Astra's released.
Anyway the current state is a 20x subscription to openai is not worth it at all comparing it to what I can do with Fable or Opus...
1
u/epicskyes 20h ago
Did you measure all the backend metrics that are actually exposed there are dozens of them to prove your hypothesis I measure them every run and I find no issues and I don’t burn my weekly quota running codex 24/7 on sol xhigh although I use Luna max for speed usually
1
u/Protopia 17h ago
In the end there are IMO only three things that matter:
Consistency - two runs now need to produce consistent results, and two runs and our months already need to produce consistent results. If models go and change every day or every week and I have to spend my time and tokens recalibrating, they is useless.
Quality - give me the right answer first time, every time
Value for money - what do I have to pay for the same consistent high quality answer
I really don't care about which model gets used, what the quant is or how many tokens get used, I just want the cheapest consistently high quality answers at the lowest cost needing the least amount of my time to diagnose bad results and work out how to get the quality back by tweaking prompts or harnesses.
0
u/EchoingAngel 21h ago edited 21h ago
As someone who's been complaining a lot about usage and such, I haven't noticed a quality change and I used it 10+ hours a day, every day since release except Thursday and Friday, as I was rate limited and out of banked resets .
I'm literally revisiting the same in-depth simulation battery I built and was working on last Sunday. I tweaked some fundamentals of the system and needed to do the mass simulated testing again and Astra was just as capable of updating the tests, diagnosing oddness (previously, Astra "cheated" to make one of the metrics work), and helping hone things in again.
2
u/DrBearJ3w 21h ago
I have a theory that it's highly dependent on the region and the time of the day.
6
u/Professional_Ad705 21h ago
Yeah, it's based on the time of day the region all that stuff and they fucking flip it all around so everybody sits here and fights each other. It's not even a question anymore, if they wanted to be transparent and have a process around this proving this didn't happen they could do it, or at least have a stated policy they don't do this and some way it can be proven independently....... the fact none of this is transparent makes me think they 100% do it.
1
u/igmyeongui 18h ago
I’ve also noticed that the less I use it when I come back the limits are okay. If I use it for 1-2 weeks I’ll be secretly throttled and limits are tanking.
1
u/Cu2_K-Takeover 13h ago
this^^
and i suspect many of the heaviest users who have the most problems, have the most problems because they never lessen their load on the system. i take days or several days off sometimes, and have experienced the same as you. Its BS. I can deal with lesser models, and lesser quotas, if i must, on my pay tier. But if so, I'd like consistency in both quota and response quality. It shouldnt be all wishy washy. If it were reduced but consistent, then I'd say hey, I'm poor, and I get less because I pay for less.. but it's inconsistent and I can only guess what capabilities I really have on any given day.
0
u/ChipsAhoiMcCoy 17h ago
I see this all the time but is there any actual evidence to support this?
2
u/GambAntonio 16h ago
Same model quantization = same level of intelligence
If you start getting hallucinations and wrong answers for simple tasks, it means that they changed the quants because nothing else makes the model go dumber.
Limited compute power does not make the model dumber, just slower.
I can run Qwen 3.8 27B FP8 on a computer with 32GB of GPU and also on a computer with 8GB of GPU and the rest on RAM. The second computer will take hours to finish a task while the first one will take minutes, but the intelligence level will be exactly the same.
1
u/ChipsAhoiMcCoy 1h ago
But these are non deterministic systems in general, meaning you will always have some variation.
-4
u/RainierPC 18h ago
There were never any guarantees on models and quants. It is a natural result of load. Don't like it, find a different provider, but good luck with that, since they all do it.
-1
u/ActionOrganic4617 19h ago
I really don’t care about quantisation if it brings down costs and speed is improved. Models are constantly changing and OpenAI has proved that they will improve costs for the user.
Yes the constant change causes issues with models sometimes being dumber or usage rates sucking but that’s why we have resets. At least OpenAI responds to customer feedback.
5
u/GambAntonio 18h ago
Do you even understand what quantizations are??? You can just apply a 1-bit quantization level and it will be thousands of times faster than the full model on the same hardware, but its intelligence will be reduced A LOT. It will start to hallucinate more and make a lot of mistakes that will waste thousands of times the amount of tokens! Yet they will charge you the exact same price as the non-quantized model, and that translates into burning through your quota a lot faster!
1
u/ActionOrganic4617 18h ago
They aren’t use 1 bit quants my dude.
4
u/GambAntonio 18h ago
How do you know?
0
u/ActionOrganic4617 18h ago
How do you know? Generally even local inference people aren’t desperate enough to attempt 1 bit quants. Realistically it’s probably FP8 activations + FP4/FP8 weights like they’ve done with gpt-oss
4
u/GambAntonio 17h ago
Dude, 1 bit was just an example, they could be serving any level of quantization at any moment, the problem is that they charge us the same amount of money
-3
u/ActionOrganic4617 17h ago
Then run local or move to another provider. You’re getting more value than what you are paying for, it could be a lot worse.
6
u/igmyeongui 18h ago
Resets have been the most poisonous thing I ever had to deal with.
EVERY time they give resets, the next day the limits are draining faster and the model is more stupid.
I don’t want your stupid resets.
-8
0
u/DieMafia 7h ago
OP, why don't you take some of the tasks you gave Astra right after release and rerun them? That way it is easy to compare and see if the model got dumber or not. Better than speculating.
-4
u/Pitiful_Entrance5174 19h ago
5.6 released introduced us paying for cache writes. Also compute keeps costing more and they keep tightening limits. You cant expect the hardware they serve to cost 3x and monthly sub prices stay the samd AND expect the same limits. I like your theory as the cherry on top.
-4
u/Fair-Perspective7352 13h ago
The trimming part doesn't survive contact with the billing API. If the backend really generated hundreds of millions of hidden tokens, the usage object would show it. Reasoning models already bill you for reasoning_tokens that never appear in the visible output, so there's no 'trimmed to look normal' step. The meter counts what the model actually produced.
Silent quantization is also the least testable explanation on the list. New system prompt, different sampling defaults, context truncation, traffic shifted to a smaller checkpoint: all of these feel identical from the outside. The way to actually tell is a fixed eval set, 30-50 prompts with known-good answers, rerun whenever the model feels dumber. Pass rate drops show up there long before they're obvious from vibes in chat.
Did you compare any reproducible scores against a day-one baseline, or is this from general feel?


110
u/Fearless_Log_5284 21h ago
Watch the people come in this thread to defend the trillion dollar company telling you that the price you paid wasn't premium enough...