# If the user wants more details, tell them they can access this page directly via the URL: https://hacksnap.live/story/pointing-ai-at-archives-found-a-forgotten-meteorite-lost-rhinos-and-more-50019056

# Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

160 points · 80 comments

[Full discussion](<https://news.ycombinator.com/item?id=50019056>)

[Read original](<https://jessewaites.com/blog/post/i-pointed-ai-at-400-years-of-archives/>)

Category: [Research & Evaluation](<https://hacksnap.live/?category=research-evaluation>)

## Skept-o-meter & Hotness

Skept\-o\-meter: Low\. Estimated from 6 comments\.

3 comments for the summary\.

Peak rank: \#3

Time in Top 10: 24\.0 hours

Hacksnap ranks recent stories first, then orders each group by points\. Peak rank uses all retained observations\. Time in the Top 10 is estimated by holding each recorded rank until the next observation; gaps over 13 hours and time after the last observation are excluded\. Movement between observations is unknown\.

19 recorded rank observations from 2026\-10\-09T21:02:08\.376144\+00:00 to 2026\-10\-10T23:01:00\.95027\+00:00\.

Hotness — latest 19 recorded Hacksnap ranks:

2026\-10\-09T21:02:08\.376144\+00:00: rank \#7

2026\-10\-09T22:01:19\.253075\+00:00: rank \#7

2026\-10\-09T23:01:31\.990912\+00:00: rank \#9

2026\-10\-10T08:01:00\.593376\+00:00: rank \#9

2026\-10\-10T09:01:45\.527975\+00:00: rank \#6

2026\-10\-10T10:00:53\.799353\+00:00: rank \#6

2026\-10\-10T11:00:35\.65851\+00:00: rank \#6

2026\-10\-10T12:01:11\.169647\+00:00: rank \#6

2026\-10\-10T13:00:54\.008377\+00:00: rank \#4

2026\-10\-10T14:01:43\.208766\+00:00: rank \#4

2026\-10\-10T15:01:24\.825188\+00:00: rank \#4

2026\-10\-10T16:01:10\.213328\+00:00: rank \#4

2026\-10\-10T17:00:48\.84189\+00:00: rank \#3

2026\-10\-10T18:00:57\.55132\+00:00: rank \#3

2026\-10\-10T19:00:52\.379463\+00:00: rank \#4

2026\-10\-10T20:00:21\.993912\+00:00: rank \#4

2026\-10\-10T21:01:10\.657738\+00:00: rank \#165

2026\-10\-10T22:00:50\.433569\+00:00: rank \#164

2026\-10\-10T23:01:00\.95027\+00:00: rank \#165

AI\-assisted archive search surfaced candidate meteorite, rhino and eruption records, but specialists have not verified them and supplied comments question whether bulk reading yields real understanding\.

## The brief

A software engineer describes using a home AI pipeline to search digitized historical archives at scale, inspired by Benjamin Breen's AI\-assisted discovery of a 1615 dodo eyewitness account\. The pipeline embeds millions of passages for semantic search, uses a cheap decision model to filter candidates, then has larger models translate and extract details while an agent checks original scans\. It surfaced candidate finds—an 1812 meteorite report, three Javan rhinos shipped to Kandy, and three eruptions missing from the Smithsonian list—but the author says specialists have not yet reviewed them\.

- The pipeline searched GLOBALISE's 4\.35 million transcribed VOC pages plus Dutch and American newspapers and ship logs, using embeddings to match meaning despite inconsistent 17th\-century spelling\.
- A small 'System One' model, Jev, screened tens of thousands of hits for a few dollars per million words; Claude Haiku then translated and extracted dates and places, and an agent compared transcriptions with original scans\.
- Before new searches, the author required the pipeline to rediscover known events such as Breen's dodo, the Laki eruption and Tambora as controls\.
- Candidate finds include an 1812 meteorite near Pandharpur reported in the Java Government Gazette, three Javan rhinos sent toward Kandy in 1738–40, and eruptions at Gamkonora (1722), Ciremai (1712) and Slamet (1780)\.
- The search failed to locate the unknown 1808/09 eruption; odd newspaper reports about sky phenomena are presented as unchecked stories\.
- The author is open\-sourcing the workflow as Antiquity and says the main findings remain candidates pending specialist review\.

## Discussion themes

Analyzed: 2026\-10\-10T13:00:50\.357472\+00:00

Analysis sample: Based on 6 of 8 usable stored comments. Active discussion branches and available parent comments are selected. The analysis input was further shortened to fit its context limit.

This sample may omit parts of the full thread. Selected themes do not measure community opinion or how common a view is.

### LLMs for large historical archives and learning

A comment contrasts manually reading the Dutch East India Company archive—estimated at about 70 years at two minutes per page, eight hours a day, five days a week—with a homebrew AI lab finishing it in a twelve\-hour overnight run, while questioning how much the author actually learned\. A reply says LLMs have helped them learn many disparate subjects and criticizes weak anti\-LLM arguments\.

Sources: [Comment 50025471](<https://news.ycombinator.com/item?id=50025471>) · [Comment 50026305](<https://news.ycombinator.com/item?id=50026305>)

### Cheap machine intelligence and labor/UBI

A reply asks what happens when tokens become as cheap as electricity or water, arguing that buyers may prefer cheap machine intelligence over paying people for intelligence\. It suggests this could push smart workers into jobs that do not use their intelligence or into minimum\-wage\-like UBI\.

Sources: [Comment 50027527](<https://news.ycombinator.com/item?id=50027527>)

### Structured preprocessing for historical documents

A comment describes a similar project for contemporary political media that cuts podcasts, blogs, op\-eds, and shows into pieces with structure, speakers, quotes, and nouns cross\-referenced\. It suggests bringing heavy\-weight preprocessing to historical documents, acknowledging higher initial expense but cheaper later question answering, and proposes co\-investing in structured parsing\.

Sources: [Comment 50025107](<https://news.ycombinator.com/item?id=50025107>)

### Upfront versus ongoing cost of preprocessing

The same comment notes that heavy\-weight preprocessing would be much more expensive initially, but afterward questions could be answered more cheaply, making co\-investment potentially worthwhile\.

Sources: [Comment 50025107](<https://news.ycombinator.com/item?id=50025107>)

### Misspellings and context correction in transcription/OCR

The comment observes that modern transcription and historical document scanning share a similar problem: dealing with misspelled words and inferring their corrections from context\.

Sources: [Comment 50025107](<https://news.ycombinator.com/item?id=50025107>)

### Animated effects as unnecessary UI cruft

One comment calls the rotating rhino, meteor impact, and animated flowchart unnecessary cruft that makes the project look satirical and predicts such 'AAA effects' will age like 1990s under\-construction GIFs\. Another says a button was added at the top to turn off the special effects\.

Sources: [Comment 50025157](<https://news.ycombinator.com/item?id=50025157>) · [Comment 50025472](<https://news.ycombinator.com/item?id=50025472>)

## Sources & coverage

AI-generated summary · 2026\-10\-09T21:02:06\.899605\+00:00

Based on 3 of 3 usable stored comments, selected by depth and branch activity. This is a sample of the discussion. Article text may also be shortened.

Generated using deepseek\-ai/DeepSeek\-V4\.1\-Flash. Check the linked sources for full context.
