TimeCapsule LLM is a unique AI model exclusively trained on texts from specific historical periods to minimize modern bias. By focusing on language and perspectives from eras like 1800-1850, this project aims to authentically recreate the worldview of the past, offering users a direct connection to historical language usage.
TimeCapsule LLM is an innovative language model designed to minimize modern biases by exclusively training on textual data from specific historical periods. This project aims to authentically capture the worldview and language reflective of its selected era rather than merely simulating a historical perspective.
Built on the foundational framework of nanoGPT by Andrej Karpathy, TimeCapsule LLM distinguishes itself by focusing on texts produced between 1800 and 1850, specifically sourced from London. The primary goal is to develop a model that inherently lacks recognition of contemporary concepts, language, and reasoning patterns that have evolved beyond the historical context.
As of July 2025, preliminary results demonstrate the model's capacity to engage with 1800's language, although complexities in sentence structuring and reasoning still require further development. For instance, when prompted with a question about a historical figure, the model produced sentences indicative of its training environment, albeit with some nonsensical phrasing.

Moving forward, the project aims to refine the model through extensive training with an expanded dataset while ensuring that all sources remain true to the historical context. The meticulous selection of texts will avoid modern interpretations, supporting the overall objective of achieving a model that convincingly mirrors the language and reasoning of its designated era.
This project emphasizes the curation and preparation of historical datasets, as well as the development of a customized tokenizer. Comprehensive training instructions for nanoGPT can be found in Andrej Karpathy's repository.
For detailed processes in gathering, preparing, and training your own historical language model, refer to the included scripts and documentation within this project.
No comments yet.
Sign in to be the first to comment.