AI Models

Talkie: 13B Language Model from Pre-1931 Texts

Researchers Nick Levine, David Duvenaud, and Alec Radford released talkie, a 13B language model trained on historical English text before 1931. The base version used 260B tokens, while the instruction-tuned model powers a chat interface. Both models carry an Apache 2.0 license, with training data free from copyright restrictions.

Neura News

Neura News

Neura Market Editorial

April 28, 20263 min read
Talkie: 13B Language Model from Pre-1931 Texts

Talkie: 13B Language Model from Pre-1931 Texts

Nick Levine, David Duvenaud, and Alec Radford, known for their work on GPT, GPT-2, and Whisper, launched a new project called talkie. This includes two models: talkie-1930-13b-base at 53.1 GB and talkie-1930-13b-it at 26.6 GB. The base model consists of a 13B language model trained on 260B tokens from historical pre-1931 English text. Alec Radford brings experience from OpenAI, where he contributed to early generative models like GPT and GPT-2, and later to Whisper for speech recognition. David Duvenaud works in machine learning research, often focusing on probabilistic models and neural networks through affiliations like the University of Toronto and the Vector Institute.

Model Specifications and Access

The instruction-tuned version, talkie-1930-13b-it, comes from a checkpoint fine-tuned on a dataset of instruction-response pairs pulled from pre-1931 reference works. This setup supports a chat interface. Users can access a demo of this chat model online. Both models operate under the Apache 2.0 license. The base model's training data falls entirely out of copyright, given the U.S. cutoff date of January 1, 1931. Expectations exist for the team to release this training data later.

Training Process and Research Goals

The project's report outlines key research aims for models like this. Simon Willison, a prominent figure in data tools and LLM commentary through his Datasette project, expressed interest in what he terms "vegan models." These are large language models trained solely on licensed or out-of-copyright data. The base talkie model fits this category. However, the chat model relies partly on modern models during fine-tuning, so it does not qualify fully.

Fine-tuning started with generating instruction-response pairs from structured historical texts. Sources included etiquette manuals, letter-writing manuals, cookbooks, dictionaries, encyclopedias, and collections of poetry and fables. The base model then trained on these in a simple chat format.

To boost instruction-following, synthetic prompts covered tasks like document summarization, direct information responses, and multi-turn conversations. Online direct preference optimization applied to rollouts from these prompts, judged by Claude Sonnet 4.6.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

A final supervised fine-tuning round used rejection-sampled multi-turn synthetic chats between Claude Opus 4.6 and talkie. This addressed remaining issues in conversation skills.

Challenges with Contamination and Future Aspirations

A primary hurdle involved preventing contamination. This meant keeping post-1931 text out of the training corpus and stopping modern LLMs from adding anachronistic knowledge during fine-tuning.

Willison recently covered Mr. Chatterbox, another vintage language model project. It faced similar needs to use modern LLMs for synthetic conversations in chat training.

The talkie developers aim to overcome this. They note that reinforcement learning with AI feedback shapes behavior in modern ways. For example, the 7B version of talkie began using listicle styles after such training. Plans call for scaling up and using the vintage base models as judges. This would create a fully bootstrapped, era-appropriate post-training pipeline.

Demo Performance

In testing the talkie demo, Willison used his standard prompt: "Generate an SVG of a pelican riding a bicycle." The model replied: "An SVG of a pelican riding a bicycle was generated in 1860. The pelican was represented as perched on the saddle, with its long bill pointed forward, and its legs astride the handles. The idea was suggested by the fact that pelicans are sometimes seen fishing on horseback, on the banks of the Rhine."

This release highlights efforts to build language models rooted in historical data, avoiding modern copyrights and influences where possible.

Related on Neura Market

More from Neura News

Industry

Monday.com Joins Tech Layoff Trend Citing AI as Factor

Monday.com announced it will lay off about 20% of its workforce, or over 600 employees, citing a restructuring tied to its AI-driven growth strategy. The Tel Aviv-based work management software company joins a growing list of major tech firms, including Amazon, Meta, and Microsoft, that have cited artificial intelligence as a factor in job cuts this year. A new Financial Times analysis shows U.S. tech companies have slashed nearly 140,000 jobs since January, with AI often cited as a reason.

Jul 26·12 min read
General

Open-weight AI mirrors Kubernetes ecosystem shift

Tobi Knaup, co-founder of Mesosphere, draws parallels between the rise of Kubernetes and the current trajectory of open-weight AI models. He argues that open-weight models are becoming a neutral substrate for innovation, attracting a global ecosystem of developers, startups, and enterprises. The piece warns against US restrictions on Chinese open-weight models, advocating instead for American leadership through open releases, procurement strategies, and standards.

Jul 25·7 min read
General

Open-weight AI mirrors Kubernetes rise, US warned on bans

The author, a Mesosphere co-founder, draws parallels between the rise of Kubernetes and the current open-weight AI ecosystem. He argues that open-weight models are becoming a neutral platform for innovation, and warns that US restrictions on Chinese open-weight models could isolate American developers from a global ecosystem. The piece urges the US to compete by releasing frontier models, using procurement to create demand, building the stack, and setting standards rather than imposing bans.

Jul 25·7 min read
Industry

Power line failure reveals AI data center grid risks and solutions

A fallen power line near Washington, DC caused over 3 gigawatts of data center load to vanish from the PJM grid in seconds, spiking voltage across the region. The event, which made lights flicker from Northern Virginia to Chicago, highlights a growing problem as AI data centers become larger and more concentrated. Experts warn that without better coordination or technology like ON.Energy's battery-backed uninterruptible power supply, such disruptions will become more frequent and severe.

Jul 25·5 min read