Dev.to · 10 min read

‘It Answers in Poetry Now’: A Week of Users Saying Their AI Got Wordier and Worse

‘It Answers in Poetry Now’: A Week of Users Saying Their AI Got Wordier and Worse

For months the loudest complaints in consumer AI have been about money and access: the meter, the rate limit, the price rise, the paid feature that quietly vanished. This week the grievance moved somewhere less obvious and more interesting — the writing. Across the big subreddits, heavy users of Claude, Gemini and ChatGPT converged on the same complaint, in almost the same words: the models have started sounding cleverer and answering worse. More verbose, more performative, more jargon, more hedging — and, underneath the flourish, less of the plain, correct answer people actually wanted. Quotes sourced from: Reddit — r/GeminiAI, r/ClaudeAI and r/OpenAI. Every quote below was read verbatim on the live thread and is listed with its username, subreddit and the thread’s date and permalink in the Sources section. We widened the window to the past week, because a two-day snapshot was thin, and we’ve deliberately included the users who disagree. As always we quote experiences, not verdicts — “the model got dumber” is one of the most over-claimed lines in AI, and we’ll say so more than once — but the specific, checkable version of the complaint is worth reading. The moan of the day: sounding smart, saying little The sharpest framing came from HPaternalPatriarchy on r/GeminiAI, in a post arguing the labs are chasing the appearance of intelligence over the real thing. The line that stuck: “It’ll easily spend around half or more of a response writing poetry while only 1 point out of 6 is actually valid or relevant to the question you asked.” — HPaternalPatriarchy, r/GeminiAI, 19 August 2026 That’s the whole complaint in one sentence: a model that fills the page confidently while the useful content thins out. The post’s thesis — that labs are “optimising for sounding intelligent instead of actually being intelligent” — is opinion, and we treat it as such. But it named a feeling a lot of people recognised this week, across more than one tool. Claude: ‘Claudish’ and the verbosity backlash The most concentrated version was on r/ClaudeAI, where Opus 5’s prose has become its own running argument. In a widely-read thread, coocoocoo8956 described “the epidemic of Claudish and Opus 5’s degrading outputs and overly verbose, hard-to-read language” — then, to their credit, walked back their own headline, noting they’d called it “dementia” when “it’s more bloat/instructional drift, my bad on the phrasing.” That self-correction is the right instinct, and rarer than it should be. Others in the thread were blunter about the impact on real work. ArcticAcademic: “the linguistic capabilities have declined drastically over the past month or so. The copy that I’m getting these days is largely unusable, despite strict restrictions and guidance. I have not changed my approach that much but Claude has obviously changed.” And ConstantKooky3329, itemising the failure modes: “word salad response to a simple simple query; tendency to overscope tasks; propensity to offer opinions when the task is to produce an output based on data. It’s definitely more rude and defensive in tone whenever you push back.” The most relatable was shanejyo, describing the loop the verbosity creates: “It’s annoying af. The past week has been ‘PLEASE REWRITE YOUR SUMMARY IN A SIMPLER LANGUAGE’ x9999.” When a chunk of your prompts are spent asking the model to say the thing it just said, but shorter, the verbosity isn’t a style preference — it’s a tax on your time. The complaint jumped subreddits, too. Back in that r/GeminiAI thread, Altruistic-Skill8667 reported the newest Claude models “drift into hardcore technical eloquent SLANG when you start talking about different professions… Sometimes it’s so bad and so cryptic that I have to look up a word or a phrase, which has never happened with other models.” InterestProof1526 put the cost bluntly: “Claude and GPT are nearly unusable for me for this reason.” When you have to look up the model’s vocabulary to use the model, the eloquence has stopped being a feature. Gemini: worse, and occasionally confidently wrong On r/GeminiAI the complaint had a harder edge, because it wasn’t only about style — it was about basic reliability. Pretend-Detail2099 asked the question that titled half the sub this week: “Has anyone noticed Gemini has gotten significantly worse recently? It seems to often misunderstand what I’m trying to say and also has started making a lot of grammatical errors.” gr1ri was more specific still: “Grammatical errors, spelling errors, wrong references and made up numbers.” Others noticed it break in oddly specific ways. LanaZ61 described Gemini bleeding one language into another: working in German on an English text, “Gemini sometimes adds random German words into the English text… Never ever had something like that with chatgpt.” It’s a small, concrete, checkable failure — exactly the kind we trust more than a sweeping ‘it got dumber’. The vivid one came from Motor-Intention4081, and it’s the kind of confidently-wrong-then-instantly-backpedal behaviour that erodes trust fastest: “I showed it some of my soldering work before, and it freaked out saying I had created a ticking time bomb, and that the job was sub par and dangerous. I told it to reassess the image and it apologized for hallucinating and that everything was fine (which it is).” An answer delivered with total confidence, then abandoned the moment it’s challenged, is worse than a hedge, because it teaches you the confidence means nothing. It’s the same trust problem we keep circling: a fluent wrong answer is harder to catch than an obvious one, which is why hallucinations remain unsolved in practice even as the prose gets slicker. ChatGPT: the quiet defection The pattern isn’t confined to one lab, and the clearest sign is people voting with their subscriptions. On r/OpenAI, Square_Secretary_944 described drifting away from a rival coding tool as its output degraded: “It started with Sol finding problems in Claude design, one or two. Then things got worse, the mistakes became bigger, Claude even owned 90% of them and they became real gaps, either in architecture or review and audit.” We can’t audit anyone’s workflow, and a switch story always flatters the tool being switched to. But a comment underneath it captured the real lesson better than any single verdict — katoptronophile, warning that the ranking is a moving target: “the performance and value ranking of these models can change suddenly.” That volatility is the actual condition now. Today’s best model is a temporary state, which is why the switching costs the vendors are relying on keep looking flimsier. The tell: rolling back to older models If one behaviour separates this week’s complaints from ordinary grumbling, it’s the rollback. When users prefer an older version of a product to the newest one, and take steps to avoid the upgrade, that’s a stronger signal than any star rating. On r/ClaudeAI, people traded methods for taming Opus 5’s output, some reporting that older releases like Fable or 4.6 read better even where the newest model benchmarks higher. It’s the same dynamic we documented when a ‘newer’ AI felt like a downgrade, and a cousin of the complaint that you paid for the big model and got served a smaller one — except this time it’s not about which model you were routed to, but about the newest one genuinely being harder to work with. Some of the theories about why got creative, and we file them as theories. One r/ClaudeAI poster, cool_architect, wondered whether Opus 5’s verbosity might be tied to the “new watermarking/SynthID feature” — a link to statistical text-watermarking that we can neither confirm nor rule out, and that Anthropic hasn’t stated. Others guessed verbosity exists to burn billable tokens, or that models are being quietly “quantised” to save compute. These are guesses about motive; we quote them as the users’ own speculation, not as fact. The fair version: capability up, prose down — and some are happy Now the counterweight, because it’s substantial and we went looking for it. Not everyone thinks their AI got worse, and the sharpest dissent was on r/ClaudeAI too. xepherys: “I’ve been using Opus 5 since it released and couldn’t be happier with it… either I’m some sort of Claude-whisperer, or most people don’t understand how to use LLMs.” The thread’s own auto-summary was honest about the split, noting the community leans frustrated but is genuinely divided. Two comments did the real work of complicating the story. Head_Leek_880 reframed the verbosity as a manageable personality rather than a defect: “Have you ever worked with a coworker who is smart but use big words and long winded? That is how I feel about opus. I don’t hate it, but I would rather let it do the work and minimize our communication.” And durable-racoon offered the most important distinction of the week — that capability and prose can move in opposite directions: “capability has increased, but prose and personality has worsened. My theory is this is due to heavier RL.” If that’s right, the models really are getting smarter and more annoying at once, and the complaint is aesthetic and ergonomic rather than about raw ability. The most useful sceptic was SaltsMoon, who named the thing everyone in these threads should keep in mind: without a controlled test, none of us actually knows. “A small replay set of old prompts may be the only way to tell whether this is a model change, prompt accretion or just a bad session.” That is the correct standard, and almost nobody meets it, us included. What to take from it, fairly Hold the caveats firmly. These are self-selecting subreddits full of power users; the contented majority rarely posts; model output is non-deterministic; and “it got nerfed” is claimed after every release, usually without evidence. None of this proves a company degraded anything on purpose, and we’ve separated the checkable complaints from the theories about motive. But notice the shape of what’s being said. It isn’t “AI is soulless” or “AI is a bubble.” It’s specific and consistent across three separate vendors: answers that are longer and denser than the question warranted, a plain response buried under performance, and enough friction that experienced users are rolling back or switching. The fixes are unglamorous, and they’re the same ones good writing has always required: Default to plain. Answer the question first, in the fewest words that are still true, and let the user ask for more — don’t make them ask for less. Don’t perform expertise. Dense jargon and rhetorical flourish read as competence to a benchmark and as noise to a human trying to get something done. Separate style from correctness. A confident tone is not a substitute for a checkable answer, and a model that backpedals the instant it’s challenged should never have sounded so sure. Let users keep what works. If a newer model reads worse for someone’s work, the option to stay on the older one isn’t nostalgia — it’s the least a paying customer should get. None of that is a revolution. It’s the difference between a tool that sounds clever and one that is actually useful — and this week, across more than one company, a lot of people felt the gap widen. Originally published at theaidownside.com — evidence-first reporting on the costs and trade-offs behind AI products.

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More AI & Machine Learning News