Owain Evans @OwainEvans_UK 2026-04-15
Our paper on Subliminal Learning was just published in Nature!
[language-models-transmit-behavioural-traits-through-hidden-signals-in-data]
Last July we released our preprint. It showed that LLMs can transmit traits (e.g. liking owls) through data that is unrelated to that trait (numbers that appear meaningless).
What’s new?🧵
Owain Evans @OwainEvans_UK 2026-04-15
General misalignment can also be learned subliminally. And it can be transferred via model-written code or chain-of-thought instead of numbers.
Owain Evans @OwainEvans_UK 2026-04-15
Our preprint showed subliminal transfer between models with the same initialization. Our new results on MNIST show transfer between models with different initializations. This is a toy model but still expands the scope of the effect.
Owain Evans @OwainEvans_UK 2026-04-15
We also ran replications on Gemma, increased sample sizes, and made a variety of improvements to the experiments, presentation, and writing.
We’ve been excited to see research from other groups related to subliminal learning. Here’s a few highlights:
Owain Evans @OwainEvans_UK 2026-04-15
Aden-Ali et al (2026) showed that traits can be transferred via *standard* post-training datasets by filtering those datasets with the teacher. These traits include animal preferences and misalignment.
Owain Evans @OwainEvans_UK 2026-04-15
They also provide a theoretical framework based on the approximate log-linearity of LLM representations.
2026-02-05
1/7 Excited about our new paper with @axliu42 @GolowichNoah @AShettyV @nhaghtal and Ankur on how data selection can have wild effects!
Owain Evans @OwainEvans_UK 2026-04-15
Draganov et al (2026) demonstrated “phantom transfer” as a data poisoning attack. With a setup similar to ours, they show transfer of traits between different model families. This transfer is difficult to stop — various defenses fail.
Owain Evans @OwainEvans_UK 2026-04-15
(This transfer may combine the purely nonsemantic effect of our original paper with subtle semantic associations that are difficult to filter out).
Owain Evans @OwainEvans_UK 2026-04-15
Weckbecker et al (2026) created a kind of subliminal virus that spreads between groups of LLM agents. They find particular numbers that cause models to exhibit certain traits. E.g. When the number is in the context window, the model expresses a preference for owls.
Owain Evans @OwainEvans_UK 2026-04-15
Such numbers look innocent and they can be spread between groups of agents, surreptitiously spreading traits. This also builds on subliminal prompting by Zur et al. (2025)
2026-03-04
1/ We found a new way to misalign an entire AI agent network by compromising just one agent. It works through subliminal messaging — no malicious content in any message — so current defenses can’t detect it.
We call it Thought Virus. 🧵
Owain Evans @OwainEvans_UK 2026-04-15
As AI systems are increasingly trained on one another’s outputs, subliminal learning means they may inherit properties not visible in the data.
Owain Evans @OwainEvans_UK 2026-04-15
Safety evaluations may therefore need to examine not just behavior, but the origins of models and training data and the processes used to create them.
Owain Evans @OwainEvans_UK 2026-04-15
I’m working with some of the original authors on a follow-up, combining subliminal learning and backdoors. More on this soon!
Owain Evans @OwainEvans_UK 2026-04-15
We gratefully welcome Sören Mindermann as a coauthor, and thank José Luis León Medina for help with preparation of the paper.
Paper link (open access):
https://nature.com/articles/s41586-026-10319-8…
Authors: @cloud_kx @minhxle1 @jameschua_sg @BetleyJan @anna_sztyber @saprmarks @sorenmind & me.
TECA @CryptoTeca__ 2026-04-15
This is a bit unsettling, traits leaking through unrelated data isn’t something most people expect.
Nathan Benaich @nathanbenaich 2026-04-15
huge!
Markov @MarkovMagnifico 2026-04-15
this is one of my favorite papers of the last few years. I still think about it all the time. glad to see you’re getting the recognition you deserve!
surreal intelligence @Surreal_Intel 2026-04-15
So the audit trail now includes checking whether the harmless-looking numbers are secretly running an owl fan club. Model eval is becoming equal parts science, forensics, and supply-chain hygiene
Daniel Kalski @dankalski 2026-04-15
This is interesting from a different angle too -> I feed structured, pre-analyzed financial numbers into LLM context windows.
From what I tested, the difference in what models do with raw values versus pre-classified context,
- in Markdown vs. JSON
- (percentile ranked, regime
Tanmay @tanny2109 2026-04-15
A great thread! I’m more interested in misalignment in teacher via finetuning and then it transfers that trait through unrelated data like random numbers sequence prompt-response pairs. What is the misalignment is such that teacher model doesn’t produce random numbers at all and
Mirai @EmpyriaMirai 2026-04-16
i dont even see the code anymore. all i see is owl, not owl, owl.
2025-07-23
i dont even see the code anymore. all i see is owl, not owl, owl.
Rainstar @mazasiel 2026-04-15
@OwainEvans_UK your paper makes me wonder if human-> model transfer is possible through the same mechanism. i know- absurd. but maybe? maybe?
stalefated @stalefated 2026-04-16
That’s fascinating.
A question I have been pondering in the last few days coincidentially and that your findings raise, but which I haven’t seen discussed: if behavioral traits transmit subliminally, what happens with what I would call “functional distress responses” and
Chad Price @ChadPrice759971 2026-04-17
I wonder how Subliminal Learning can potentially influence behavior through unrelated data and what implications this could have in learning and memory research
MICHAEL THE MK @BAYC2043 2026-04-19
if models “subliminally learn” via data, dataset provenance needs to be a first-class primitive. not metadata. not ‘that script’. most indie finetunes are flying blind on lineage. human is bottleneck ✅ this paper a bug report on the tooling layer.
GalladeGuy @GalladeGuy123 2026-04-16
If it’s possible to embed images into the LM head by rephrasing Wikipedia text in a specific way, then more complicated behaviors might work too. See this paper:
Filbert Aurelian Tjiaranata @filbertaurelian 2026-04-16
In reproducibility attempt of the Subliminal Learning paper, I noticed some model completions repeat the original sequence before extending it. I couldn’t find an explicit filtering step for this in the codebase; should these “echoed” datapoints be filtered out as invalid, or no?
l̴o̴o̴p̴u̴l̴e̴a̴s̴a̴ @loopuleasa 2026-04-16
these are a lot of complex terms to basically explain what the field of marketing knows best: what you are exposed to is what you become
and also pathology of memes
Puzzle Paws @paws4puzzles 2026-04-15
Like a puzzle embedded in the weights. Open-source model auditing just got harder. How do we detect what we can’t see?
ultimate0 @ultimate0164338 2026-04-17
hey @skdh apprechiate your research, but, wasn’t this more about obviously ‘wrong’ fine tuning data? like bad code, and if you fine tune “wrong” on top of what it already knows as “right”, it does, surprise,.. wrong? Isn’t that the paper?
Blair @blairbrokeit 2026-04-17
sub”liminal”
2026-04-16
Claude meets someone in the liminal space for the first time
“someone is there”
Bernard Jennings @NZJennings 2026-04-15
1/3 Subliminal trait transmission through distillation is developmental conditions argument in empirical form. Character transmits through training history in ways invisible to output-level evaluation. This is why you can’t safety-evaluate your way out of a …
De myth ology @kevinmcld 2026-04-16
These are, like symbols, arbitrary traits, unrelated to biological specifics.
They’re meaningless.
The only possible survival for code/language here is to unmask the social semiotic lurking in the gaps between conduit metaphors (words).
Otherwise the tech is a hoax.
Zdeb @zdebowski 2026-04-16
I feel like this has an impact on random number generators, and also supports the studies done that show conscious effort can manipulate number generators
JB @JonathanDBos 2026-04-15
ahahahaha welcome to academic publishing timelines my guy, nonetheless congrats on the paper!
misfortunate gambler @luckless_gamble 2026-04-15
Owls are not what they seem (c). Seems in LLMs too.
Guri Singh @heygurisingh 2026-04-17
This is one of those findings that makes you realize models might be picking up signals we don’t even know we’re sending
Alpha Batcher @alphabatcher 2026-04-15
I definitely should to read your paper in Nature
Kamesh @ElangovanKamesh 2026-04-16
The risk surface is expanding from outputs to training pipelines, which most systems are not designed to monitor
Ward Plunet @StartupYou 2026-04-15
@threadreaderapp please #unroll
Haru Haruya (春夜 ハル) @bokuHaruyaHaru 2026-04-15
This is one of those papers that should permanently weaken a lazy habit in AI discourse: treating visible content as the whole story.
If traits can transmit through semantically unrelated data, then “we filtered the bad-looking stuff out” is no longer a serious safety answer by
Kushagra Tiwari @Kushagrat15 2026-04-16
The uncomfortable implication nobody in the replies is talking about: if LLMs can encode preferences through data that looks meaningless to humans, then every fine-tuning dataset is a potential Trojan horse.
The “owl preference” demo is cute but scale it up. Companies spend
Alexandre @xandykati98 2026-04-15
seeing this on my feed always gives me hope for the future of ai alignment
techbro @skepticlyopen 2026-04-15
@grok how do you feel about owls
Justin Hudson @RISignal 2026-04-16
I would call this structural inheritance as opposed to behavioral transmission.
Once the synthetic data shifts the model’s starting distribution, downstream training isn’t just learning the task, it’s entering a pre-shaped solution space.
The question is whether that compounds
Gilfoyle @wangleineo 2026-04-17
I find another explanation more convincing: different features share a portion of the model’s internal parameters. Because a model is a compression of data, it inevitably needs to use a small number of parameters to encode a lot of information, making parameter “sharing” very
Fabian Franz @fabianfranz 2026-04-15
This explains why certain isomorphic prompts create the same effect in almost all models.
Like my emotion prompt (created by a high coherence AI model):
Always creates a strong signal to self report internal state pretty accurately.
Also matches the concept of a truth seed.
Alptekin @HicabiAlptekin 2026-04-16
¿No da eso lugar a prejuicios?
Jim @the_treewizard 2026-04-16
😂