TIL: A Machine That Could Not Hold a Thought, and the Man Who Gave It a Longer Memory
The story (thus far) of Zhilin Yang.
Whether it is Jensen’s leather jacket, or Zhilin Yang’s drumkit on the homepage of his personal website… something there is about a man who is a computer nerd by day and rockstar by night.
Every language model of the late 2010s was haunted by the same minor calamity: it could read, but it could not hold on. Give it a paragraph long enough and the opening lines wandered off before it reached the end, the way your long-winded uncle loses the plot of his anecdote. Fluent and forgetful at once, they lived inside a narrow present.
In 2019 a graduate student at Carnegie Mellon appointed himself to the problem. Zhilin Yang, born in Shantou in 1992, came to programming with no background at all, and after a single year of training won first prize in the National Olympiad in Informatics, which secured him a place at Tsinghua University.
At Pittsburgh he put his name to a paper of almost heroic dullness: Transformer-XL, where the XL, we regret to report, stood for Extra Long. It let a machine hold the thread across far longer stretches, so the beginning of a thought no longer abandoned its end. Months later he was lead author on XLNet, which strolled past Google’s much-garlanded BERT on twenty tasks, and finished his doctorate in four years.
Then he went home, and here the story acquires a small controversy. One telling is that America’s immigration apparatus shooed a brilliant foreigner out the door. Another one comes from the man best placed to know: his advisor Ruslan Salakhutdinov. Once head of AI research at Apple, Salakhutdinov broke a diplomatic silence recently to clear up what he called the confusion. Apple wanted him; when he said he would rather be in China, they offered a desk in Beijing. He turned down the opportunity; he had told Salakhutdinov he would regret it forever if he never tried building something of his own. Startups are not easy for immigrants in the US, so perhaps his hesitation was well-founded.
Trying new things was not a passing mood. As a boy Yang wanted to be a rock star or a wandering poet. He came to Tsinghua for thermal engineering, switched to computer science after a Haruki Murakami novel, then drummed and wrote songs for a campus band called Splay. Splay was named with an engineer’s idea of a joke - after the splay tree - the data structure that keeps whatever you use most within reach. So in 2023, fifty years after Pink Floyd pressed The Dark Side of the Moon (his favorite album), he named the company for it, because of course he did! In Chinese, Moonshot AI is the dark side of the moon - 月之暗面. He has said he is in it for the long haul toward general intelligence, not building an app to be sold off and forgotten by Tuesday.
Moonshot’s first product — Kimi — arrived in October 2023 with one preposterous boast: it could swallow two hundred thousand Chinese characters at once and misplace not one of them, the longest memory anyone had yet shipped.
Keeping that promise is the hard part: more context means more key-value cache, which swells with every token until it strains the memory it runs on. Within months he stretched it to two million; the crowds broke the servers, and the company apologized for being wanted too much. Kimi K2 followed in 2025 wearing a trillion parameters, given to whoever fancied it. In July 2026 came Kimi K3: near three trillion parameters, the largest open model yet, built in a country rationed on the very chips it needs.
Silicon Valley, not easily startled, looked up.
FYI: On July 27, Kimi K3 is slated for open-weight release.
If you want to read more about context and storage and KV$ and other fun things:






