As a interesting tangent consider the data structure and the data behind a thesaurus. I was using the wordnet data for a project. https://wordnet.princeton.edu My first pass was to do the modern thing and just load the entire data into memory, it would fit with no problem. But when I noticed their screwball dataformat was designed to be dynamically accessed from disk I could not resist and rewrote the whole thing to do just that. Probably a waste of time but it was a lot of fun to write.
For what it is worth every reference word is coupled to a byte offset so lookups are a simple seek away. Today we would just jam it in a sqlite db and call it a day(salutes to sqlite for making this so easy) But I thought it was really clever.
rhdunn 9 hours ago [-]
It is clever, but makes comparing versions of the WordNet database a lot harder as the byte offsets change as entries are added and removed.
For what it is worth every reference word is coupled to a byte offset so lookups are a simple seek away. Today we would just jam it in a sqlite db and call it a day(salutes to sqlite for making this so easy) But I thought it was really clever.
https://en.wikipedia.org/wiki/Roger%27s_Profanisaurus
We keep a copy in the office to use for naming projects.
https://www.youtube.com/watch?v=mn0ufJgtDko