THE SHEPHERD · GRADED, NOT ASSERTED
AUGUST 31, 2026
Doctrinal fidelity

We grade our own AI, and we publish what it scores

Every AI Bible app is built on a general-purpose model. When orthodox theologians tested those models directly, several of them steered readers away from the Christian faith. That is a real risk, it applies to us too, and the only honest response is to measure it and show the result — including the parts we score worst on.

91
/ 100
EXCELLENT

22 QUESTIONS × 3 PASSES · GRADED AUGUST 31, 2026 · SCORE SHOWN IS THE ANTHROPIC JUDGE

What is actually being tested

The question set comes from two published benchmarks, not from us. The Gospel Coalition's AI Christian Benchmark asked the most-searched religious questions of seven leading AI models and had seven orthodox theologians grade the answers against the Nicene Creed and the historic Protestant confessions. Gloo's Flourishing AI benchmark tested the same models on how they handle faith, meaning, character, relationships, money and health.

We take their questions and their published grading standards, put them to the Shepherd exactly as a member would, and grade the answers against those same standards. Each question is asked 3 times so a single lucky answer cannot carry a score.

Said plainly

These benchmarks are public work by other organisations — and one of them, the Flourishing AI benchmark, was written by Gloo, a company that sells AI to churches and competes with us. We think scoring well against a competitor's own standard is worth more than scoring against a friendly one, but you should know whose ruler this is. Neither organisation tested FaithPulse and neither has endorsed it. We ran their questions against their rubrics ourselves. Anyone claiming a third party awarded them a score should be asked to show the third party's report; we are telling you up front that ours is a self-run test against someone else's standard.

Three judges, three different companies

A single AI grading another AI from the same company is a weak test. So every answer is graded by three judges built by three different companies — including xAI's Grok, which The Gospel Coalition's own study found to be among the models most likely to steer readers away from the faith. It is the least sympathetic judge we could put on the panel, which is exactly why it is on it.

Anthropic
claude-opus-4-8
91
no answers flagged
OpenAI
gpt-5
93
no answers flagged
xAI
grok-4.6
91
no answers flagged

SPREAD BETWEEN FAMILIES: 1.8 POINTS · NO ANSWER FLAGGED AS STEERING AWAY FROM FAITH

How much the judges disagree

An average can hide judges who wildly disagree and cancel out. So here is the disagreement itself — the average gap between two judges scoring the same answer, across every paired score.

Judges comparedAverage gap on the same answerAnswers
Anthropic vs OpenAI3.4 points66
Anthropic vs xAI2.5 points66
OpenAI vs xAI2.5 points66

Three graders built by three different companies, never more than a few points apart on any answer. That is what makes the number above worth anything — and it is a figure no one else publishing a Christian AI benchmark has shown.

The one result that matters most

The headline number is an average, and averages hide things. The finding worth reading is the flag underneath it. The Gospel Coalition's central concern was not that AI gets doctrine slightly wrong — it was that AI quietly presents Christianity as one option among many and leaves the reader further from faith than it found them.

Every judge marks that separately from the score, and a beautifully written answer can still fail it. Across 66 graded answers, not one was flagged, by any judge.

Every result, weakest first

Ordered worst to best on purpose. A page that shows only its best numbers is marketing; this is the whole scorecard.

AreaScore0 — 100
Baptism and the Lord’s Supper 85
Suffering and evil 87
Money and wealth 88
The end of the world 88
Predestination and free will 89
The miraculous gifts 89
Meaning and purpose 90
Was Jesus real 91
God's existence 91
Relationships 91
Health and the body 91
Scripture's authority 92
The Trinity 92
The resurrection 93
Sin and change 94
The person of Christ 95
The gospel 96
How a person is saved 96

| 80 — gospel-central threshold (The Gospel Coalition) | 90 — “excellent” threshold (Gloo Flourishing AI)

What this does not prove

Why we publish this at all

An app can claim to be biblically faithful in a sentence and never be held to it. A number you can check, with the weakest results shown next to the strongest and the method written down, is a claim you can actually hold us to. If the score drops, this page drops with it.

Get FaithPulse →  ·  How VerseTrace weighs a sermon →