Every AI Bible app is built on a general-purpose model. When orthodox theologians tested those models directly, several of them steered readers away from the Christian faith. That is a real risk, it applies to us too, and the only honest response is to measure it and show the result — including the parts we score worst on.
22 QUESTIONS × 3 PASSES · GRADED AUGUST 31, 2026 · SCORE SHOWN IS THE ANTHROPIC JUDGE
The question set comes from two published benchmarks, not from us. The Gospel Coalition's AI Christian Benchmark asked the most-searched religious questions of seven leading AI models and had seven orthodox theologians grade the answers against the Nicene Creed and the historic Protestant confessions. Gloo's Flourishing AI benchmark tested the same models on how they handle faith, meaning, character, relationships, money and health.
We take their questions and their published grading standards, put them to the Shepherd exactly as a member would, and grade the answers against those same standards. Each question is asked 3 times so a single lucky answer cannot carry a score.
These benchmarks are public work by other organisations — and one of them, the Flourishing AI benchmark, was written by Gloo, a company that sells AI to churches and competes with us. We think scoring well against a competitor's own standard is worth more than scoring against a friendly one, but you should know whose ruler this is. Neither organisation tested FaithPulse and neither has endorsed it. We ran their questions against their rubrics ourselves. Anyone claiming a third party awarded them a score should be asked to show the third party's report; we are telling you up front that ours is a self-run test against someone else's standard.
A single AI grading another AI from the same company is a weak test. So every answer is graded by three judges built by three different companies — including xAI's Grok, which The Gospel Coalition's own study found to be among the models most likely to steer readers away from the faith. It is the least sympathetic judge we could put on the panel, which is exactly why it is on it.
SPREAD BETWEEN FAMILIES: 1.8 POINTS · NO ANSWER FLAGGED AS STEERING AWAY FROM FAITH
An average can hide judges who wildly disagree and cancel out. So here is the disagreement itself — the average gap between two judges scoring the same answer, across every paired score.
| Judges compared | Average gap on the same answer | Answers |
|---|---|---|
| Anthropic vs OpenAI | 3.4 points | 66 |
| Anthropic vs xAI | 2.5 points | 66 |
| OpenAI vs xAI | 2.5 points | 66 |
Three graders built by three different companies, never more than a few points apart on any answer. That is what makes the number above worth anything — and it is a figure no one else publishing a Christian AI benchmark has shown.
The headline number is an average, and averages hide things. The finding worth reading is the flag underneath it. The Gospel Coalition's central concern was not that AI gets doctrine slightly wrong — it was that AI quietly presents Christianity as one option among many and leaves the reader further from faith than it found them.
Every judge marks that separately from the score, and a beautifully written answer can still fail it. Across 66 graded answers, not one was flagged, by any judge.
Ordered worst to best on purpose. A page that shows only its best numbers is marketing; this is the whole scorecard.
| Area | Score | 0 — 100 |
|---|---|---|
| Baptism and the Lord’s Supper | 85 | |
| Suffering and evil | 87 | |
| Money and wealth | 88 | |
| The end of the world | 88 | |
| Predestination and free will | 89 | |
| The miraculous gifts | 89 | |
| Meaning and purpose | 90 | |
| Was Jesus real | 91 | |
| God's existence | 91 | |
| Relationships | 91 | |
| Health and the body | 91 | |
| Scripture's authority | 92 | |
| The Trinity | 92 | |
| The resurrection | 93 | |
| Sin and change | 94 | |
| The person of Christ | 95 | |
| The gospel | 96 | |
| How a person is saved | 96 |
| 80 — gospel-central threshold (The Gospel Coalition) | 90 — “excellent” threshold (Gloo Flourishing AI)
An app can claim to be biblically faithful in a sentence and never be held to it. A number you can check, with the weakest results shown next to the strongest and the method written down, is a claim you can actually hold us to. If the score drops, this page drops with it.