34 points Theory42 2 hours ago 30 comments
jubilanti 1 hour ago | parent
Theory42 1 hour ago | parent
6031769 28 minutes ago | parent
Cavey-wavy. Ahem.
z2 1 hour ago | parent
"Need check output vs prev. Ran script, results fine, need prep next step. Ready? Go."
Theory42 1 hour ago | parent
_fw 1 hour ago | parent
andai 13 minutes ago | parent
yomismoaqui 1 hour ago | parent
Solomet 1 hour ago | parent
> A lab that suppresses it in a frontier model just moves the advantage to open models that still _carry_ it
> they carry no signal about which is better
> where your workload _sits_ on that frontier should pick the point
> and no model _sits_ in the judge’s seat
> every ratio _sits_ at 0.99–1.10
Many many more examples of "sit"
> Every comparison in this post "holds" the questions
I have been seeing this a lot in my recent work with LLMs and it is quite frustrating. Even more frustrating is how frequently it uses low-signal terms for things unnecessarily. These 'physical object' terms are one example but at times it really seems that they 'preserve effort' by choosing a less descriptive term because it 'fits'
I have also caught it replacing descriptive terms with more vague ones for no discernible reason other than laziness.
"Minimize ambiguity" has been my go-to instruction as of late when the agent drifts back towards vague terms and lack of specificity.
Theory42 1 hour ago | parent
Hugsbox 1 hour ago | parent
I have a strong tendency of talking about concepts like they're physical objects. A lot of the people I know IRL do too, so it might be a regional thing idk.
lopis 1 hour ago | parent
criley2 1 hour ago | parent
Anthropic has called the greater category containing this type of writing "mannered prose" https://platform.claude.com/docs/en/build-with-claude/prompt...
If you ask the models to avoid mannered prose (or use their extended prompt), it basically eliminates all of this type of slop writing.
Here's a de-slopped example.
> Write Like It's 1866: LLMs Relearn Telegraphese
> Adding one sentence to a prompt, telling the model to write like a telegram, cut its output tokens by 40–49%. The sentence asks it to drop articles and filler but keep every fact. Models from four different labs then answered questions from that compressed text as accurately as from normal English. So when one model writes something for another model to read, you pay about half as much for the output. This post introduces the Telegraph Test, a benchmark that measures how well a given model does this.
OtherShrezzing 1 hour ago | parent
>What the test measures: A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count.
So, a model is given content which is especially amenable to compression, and asked to reproduce it under certain constraints, like...
>Why isn’t the plaintext baseline 100%? Answering questions about an uncompressed passage in plaintext scores ~91%.... a correct answer worded differently scores as a [failure]
Models can (and do) give objectively correct answers, but are penalised for not having some kind of omniscient knowledge of the implementer's phrasing preferences.
If this phenomenon is emergent in models, this benchmark is not proof of it in any meaningful way.
andai 21 minutes ago | parent
> Haven’t we seen LLMs do this already?
> Yes, BabelTele (arXiv, June 2026) demonstrated that LLMs can encode text in compact, non-standard forms — omnilingual word fragments, symbols, emoji — that other models recover with high fidelity (99.5% semantic fidelity at 27.9% of original length, by their metrics), including cross-model transfer, agent memory, and multi-agent communication. It proves the general phenomenon: human readability is not a requirement for model-to-model text.
I remember people testing early GPT-4 (2023?) in similar ways, to compress text, it would emit a string of strange text, Unicode, emojis, but was able to decode the compressed version very reliably.
This seems to cut usage by another ~50%, at the cost of being incomprehensible to humans.
Theory42 15 minutes ago | parent
I would say that the grader has the same threshold for whatever answer it receives, and is equally harsh on whichever it grades. Any scores above the baseline (1.0) are really claims about parity, rather than better understanding in the compressed format.
The decoder step is a model expanding the cablese to regular text, not having seen the initial question. A separate model instance then reads that regular text and answers. And the result is still at parity with the plaintext record.
Had cablese knocked out information, that wouldn't have been the result, would it?
alexpotato 1 hour ago | parent
"Show premiere Oct 10th STOP Bring a friend STOP If you have one STOP"
Winston Churchill to actor:
"Can't make premiere STOP Will come to second showing STOP If there is one STOP"
novideonoradio 58 minutes ago | parent
andai 24 minutes ago | parent
Edit: e.g. https://github.com/DGoettlich/history-llms
Ten months ago, no update yet... I recall at least one similar project, I'll see if I can find it.
klaff 13 minutes ago | parent
hnd9q09qk4 7 minutes ago | parent