Reading difficulty will always be subjective, based mostly on what we've already read and mastered.
But books do have objective qualities that can help us rank them along a reading difficulty spectrum.
I made this visual atlas to help figure out what works have broadly similar difficulty levels, which helps us find a good next read that's a doable step up the ladder but doesn't throw us in over our heads.
I set Familia Romana chapter one as 1.00 on a 0.00 to 10.10 scale, with every other work floating freely in comparison to it.
The metrics used are:
- Surface: Sentence length, vocabulary variety (type-token ratio), word length
- Lexical: Rare / very-rare / unseen vocabulary fraction, proper-name burden
- Morphosyntactic: Finite-verb density, nonfinite constructions, subordination, case-form diversity
- Progression: Entry burden, sustained burden, late-work burden, peak burden
- Dialogue-specific: Dialogue density, ellipsis rate (short verbless utterances), speaker-turn rate
I had this data because I've been building contextual glossing for Librorum Capsula, and had to have detailed data about each sentence. From that base, I could create rankings for entire books.
Right now I've only got a bit over 100 works of various types and lengths in the Atlas, but I expect it will become more and more useful as I build out Librorum Capsula's library.
Please let me know if you find this useful.