Skip to content
Alle Einblicke
  • Trust
  • News Systems
  • Product Architecture

Scoring news trustworthiness: designing a system that doesn't claim to be neutral

A trust label is an editorial decision encoded in software, not a measurement. Every design choice downstream of that fact — from data model to what you show the reader — depends on accepting it first.

Veröffentlicht

9 Min. Lesezeit

When a system tells a reader "this story is reliable" or "this source should be read with caution," it is making a judgment, not reporting a physical quantity. You can measure the temperature of a room. You cannot measure the trustworthiness of a news source in the same sense. Every system that attaches a trust label to a story is making a stack of editorial decisions behind that label: what counts as a signal of reliability, what counts as a signal against it, how much a source's history should weigh against its most recent story, and who gets to change that weighting. Encoding those decisions in software doesn't make them go away — it makes them silent and durable instead of debatable and visible. The real design problem isn't how to be "objective." It's how to be honest about the fact that you aren't, and to build the system so that position stays reviewable, explainable and correctable.

The label is an editorial position, not a measurement

The most common mistake is treating a "trust score" like any other computed field — something derived from a handful of signals that collapses into a single number between zero and a hundred. The problem isn't that this is hard to implement. The problem is that the number itself makes a claim the system can't back up. A single score creates the illusion that all the dimensions of trustworthiness — factual accuracy, framing bias, correction history, source transparency, speed of publication versus verification — sit on one comparable axis. They don't. A source can be accurate on raw facts and heavily biased in framing. Another can be new and have made no mistakes yet, which is not the same thing as having earned trust. Compressing these into one number quietly transfers a decision about which dimension matters most from the editorial side to the engineering side, without anyone having made that choice on purpose.

The better approach is to build the label as a declared position from day one, not a calculation. That means documenting, as seriously as you'd document a database schema, which criteria this particular system weights, why, and who owns that choice. If your team decides that correction speed matters more than publishing volume, that decision needs to live somewhere the next engineer — and a curious reader — can actually find it.

Source attribution has to be a first-class object, not a string

A more common architectural mistake is storing "source" as a text field on the story row — a name or a URL. This works today and traps you six months later, because a source has its own history: its name changes, its ownership changes, its record of corrections accumulates, and its trust ranking needs to evolve over time rather than being recomputed from scratch on every story. If the source is a string, none of that is possible — you just have a label glued to one story, disconnected from the label on that same source's previous story.

A system like Arious News, which attaches a source/trust status to every story, needs to model the source itself as an independent entity for exactly this reason: with its own identifier, its own history, and its relationship to every story it has produced. That means a source's trust ranking can be updated independently of any single story, its trajectory over time is visible, and — the part that matters most — a correction can be attributed to the correct source, not lost in a row in the stories table.

What automation can decide, and what it can't

Language models and statistical classifiers are good at detecting measurable patterns: does this text make unverifiable claims? Has this source published corrections repeatedly before? Does the language differ from typical fact-first reporting style? Those are useful signals and can reasonably be computed automatically. What automation cannot do is decide the relative importance of those signals — and, more importantly, decide the borderline cases where no signal is clear. A story that is factually accurate but timed specifically to create a misleading impression requires a human judgment call that no classifier can make on its own, because the thing being judged isn't in the text at all.

A useful working rule: automation produces signal, a human decides the threshold. If your system lets a model attach the final label to a story with no human review checkpoint anywhere in the pipeline, you have effectively delegated an editorial decision to something that cannot explain why it made that call and cannot be held accountable for it.

Corrections are a first-class state, not a deletion

When a source's or a story's trust label turns out to be wrong — and over enough time, it will — the temptation is to edit the record and move on. Resist that. A correction should be its own object in your data model: what was originally said, what was actually true, when it was corrected, and why. This matters for transparency with the reader, but it matters just as much for the system itself: if corrections aren't recorded as events, you have no reliable signal for "how often has this source been wrong, and how did it respond" — and that signal is one of the few things about trust that's actually worth trusting.

The data model should store a correction as an event in time, not as an overwrite of the current value. The prior version of the story, along with the correction's timestamp, needs to remain retrievable.

Who gets to change the weights

Every trust-scoring system eventually reaches this question: someone wants to change the weight of one criterion — deciding, say, that correction speed should count for twice as much in the final score as it did before. From the code's point of view, this change is just a number in a configuration file. From the point of view of its impact, it's an editorial decision that can reshuffle the ranking of dozens of sources overnight. If the path for changing these weights is the same path as any other code deployment — a pull request, an approval, a deploy — then an editorial decision gets made with the same amount of scrutiny as a button color change, and almost nobody reviews it as carefully as it deserves.

The fix isn't to make changing the weights harder; it's to make it traceable and owned. Every change to weighting criteria should be logged: what changed, who decided it, and what the reasoning was — the same way a correction to a story gets logged. And it should be clear whose organizational role this decision belongs to: the engineer who writes the code isn't necessarily the person who should decide which criterion matters more. Keeping those two roles distinct — even when, in practice, one person does both — forces every weight change to be seen consciously as an editorial decision, not just another commit.

What you show the reader versus what you only compute

The last decision, maybe the most important one, is how much of the internal logic to expose. The common temptation is to surface every computed signal in the interface, because "transparency" sounds like the right instinct. But showing a reader who just wants to know "should I trust this story" a dozen sub-scores isn't transparency — it's cognitive overhead that obscures the actual decision the system made.

What you show the reader should be small, stable and legible: the source's status, and where one exists, a summarized reason for that status. What you only compute — raw model scores, internal weights, experimental signals — belongs in logs and internal tooling, where an editorial team can review it without asking the reader to interpret it. The gap between these two layers is itself a design decision, and it needs to be made on purpose, not left as a side effect of whatever fields happened to land in an API response.

A good trust system doesn't claim to be neutral. It claims to know its own position, to have documented it, to let its mistakes be seen and corrected, and to draw a clear line between what it owes the reader and what is only its own internal tooling. That's exactly what a good editor does too — it's just written in code this time.

Etwas zu bauen?

Erzählen Sie uns, woran Sie arbeiten. Wir sagen Ihnen ehrlich, ob wir das richtige Team dafür sind.

Gespräch beginnen

oder schreiben Sie uns an hello@larsima.com