We used language models to screen 26 nonfiction books. They agreed on the broad winners, disagreed on the details, and helped us choose two finalists for a human trial.
Worth pinning down what the five models are actually scoring. A book can read as warm and conciliatory and still not move a single reader, while a pricklier one moves plenty. How bridging the text sounds isn’t the same as what it does to the person who finishes it. Sounds like the RCT is the part that tests that, and the model scores are a way to pick candidates, not proof they work.
Yes, this is correct: the RCT will be the part that tests whether these books actually move the reader. We won't know for sure until we finish the next phase of the study. This is just to pick the most promising candidates.
Worth pinning down what the five models are actually scoring. A book can read as warm and conciliatory and still not move a single reader, while a pricklier one moves plenty. How bridging the text sounds isn’t the same as what it does to the person who finishes it. Sounds like the RCT is the part that tests that, and the model scores are a way to pick candidates, not proof they work.
Yes, this is correct: the RCT will be the part that tests whether these books actually move the reader. We won't know for sure until we finish the next phase of the study. This is just to pick the most promising candidates.
Looks like I have some reading to do!
Jay, I heard you speak at Summit on Tuesday, you were phenomenal! I’m excited to read your book and follow along here :)