| 
View
 

Evidence Scores

Page history last edited by Mike 1 week, 1 day ago

Home > Page Design > Algorithms > Evidence Scores

Evidence Scores: When Arguments Break Under Pressure

The problem: A link to a peer-reviewed study with disclosed methodology looks exactly like a link to a government press release. Same blue underline. Same visual weight. Your brain knows one is more reliable. The interface doesn't care.

The solution: The ISE doesn't ask "who said it?" It asks "what's the argument, and does it break when you pull on it?" Evidence Scores make argument quality visible, and credentials don't get you out of making the argument.

Division of labor: what evidence is, and the Tier 1-4 source classifications, live on the canonical Evidence page. This page is the scoring side: how a piece of evidence earns or loses its impact. All numbers on this page are [bracketed] illustrations; the live scoring engine is in development.

The Formula

Every piece of evidence gets scored on two independent dimensions:

Evidence Impact = Quality Score × Linkage Score

Quality asks: how well does the methodology behind this evidence hold up when challenged? Linkage asks: does it actually prove the specific claim being made, or just point in the same general direction? Multiplying them together prevents two common failure modes: strong methodology applied to an irrelevant question, and a directly relevant claim backed by nothing but anecdote. Either factor near zero drags the impact near zero, which is exactly right.

Why Credentials Don't Determine the Score

Institutions lie. The government fabricated the Gulf of Tonkin incident to justify Vietnam. Purdue Pharma claimed OxyContin was less than 1% addictive, leaning on a single paragraph in a 1980 letter rather than a study. The tobacco industry funded research denying the cancer link for decades. In each case, the dissenters who got it right had less institutional authority than the experts who got it wrong.

This isn't an argument against expertise. NASA often produces better work than a random blog, and the ISE framework is built to show that. But not because NASA has authority. Because NASA often uses better methodology: transparent measurement protocols, redundant verification, published data others can check, and replication across different instruments. Those are the things that earn a high Quality Score. Being NASA doesn't.

When NASA failed, with Challenger and Columbia, it was because institutional pressure overrode good methodology. The engineers making valid arguments about the O-rings had sound reasoning and were outranked. A system that scored the argument instead of the org chart would have elevated them. That is the whole point: it doesn't care about org charts.

What "Quality" Actually Means

Rather than a credential hierarchy, the ISE tracks four patterns that consistently produce arguments that survive scrutiny:

PatternWhy it survives scrutinyHow credentials fake it (and get caught)
Transparent measurement with controls You showed your methodology, controlled for alternative explanations, and made your data available. When challenged, you can defend each step with specifics. Fake: "We're experts, trust our model" without showing the model. Caught: "Release your code" exposes assumptions you can't defend.
Replication across contexts Multiple independent groups using different methods arrive at similar conclusions. Hard to fake because it requires coordinated fabrication. Fake: Citing 10 studies from the same network that all cite each other. Caught: "These aren't independent, they share funding and authors."
Falsifiable predictions You said "if X is true, we should observe Y." Then Y happened. Reality validated the argument. Fake: Vague predictions retrofitted to any outcome. Caught: "Your prediction was so fuzzy it couldn't be wrong."
Explicit assumptions You stated your assumptions clearly so others can challenge them. "We assume A, B, C. If you disagree with C, here's how that changes the conclusion." Fake: Hiding assumptions in jargon or calling them "standard in the field." Caught: "That 'standard assumption' is doing all the work, and it isn't justified."

Notice what's absent from that table: journals, degrees, institutions, funding sources. Those might correlate with quality in some domains. They don't cause it, and they don't protect against its absence.

Tiers Are Priors, Not Verdicts

A careful reader will notice an apparent contradiction: this page says credentials don't determine scores, and yet every evidence ledger on this wiki stamps sources Tier 1 through Tier 4, with peer-reviewed work at the top. Isn't that a credential hierarchy wearing a number?

No, and the distinction matters. A tier encodes a methodology track record, not an authority ranking: peer-reviewed work is Tier 1 because that pipeline usually forces the four patterns above (disclosed methods, review, replication pressure), not because professors outrank bloggers. So the tier sets the starting prior for a Quality Score, and the challenge network moves the score from there, in either direction. A Tier 1 source can sink: Purdue's "less than 1% addictive" claim wore a peer-reviewed journal's letterhead, and the challenge that it was a five-sentence letter, not a study, correctly demolished its quality. A Tier 4 source can rise: an anecdote that survives verification and makes a falsifiable prediction that lands earns quality the tier never promised.

This is the same design used everywhere on the ISE: Maslow validity bands are priors the argument tree can override, never a fixed lookup table. Priors make the system efficient. Overridability keeps it honest. A prior that couldn't be overridden would just be authority with extra steps.

How Scoring Actually Works: Arguments All the Way Down

When someone submits evidence, they're not just dropping a link. They're making claims about why that evidence is reliable. Those claims form an argument network that gets evaluated the same way every other argument does.

Here's a concrete example. A government report claims "Policy X will create Y jobs," based on an input-output economic model using Bureau of Labor Statistics multipliers with an assumed Z% implementation rate.

Three challenges come in. A grad student points out the multipliers are from 2010 data and labor markets have shifted significantly since then, and shows more recent studies with lower multipliers. A credentialed think tank responds: "This model is standard in the field, used by top economists." Your uncle the accountant notices the Z% implementation rate assumes mandatory participation, but the actual legal text makes it voluntary.

The grad student's challenge is methodologically valid and survives scrutiny. The think tank's defense is an appeal to authority, which is itself an argument, scored like any other, and it scores poorly here because it never addresses the outdated multipliers. Your uncle caught a genuine assumption error. The score that emerges, using illustrative numbers: Quality around [0.40], Linkage around [0.70], Evidence Impact around [0.28]. The government's credentials didn't protect the weak methodology. Your uncle's lack of credentials didn't stop his valid challenge from counting.

Contributors do build track records over time: people whose challenges keep surviving scrutiny become worth reading, and people whose challenges keep collapsing become skimmable. But a track record is provenance for readers, never an input to any score, because who said it does not change what it's worth. The moment reputation started weighting scores, the system would be rebuilding the org chart it exists to replace.

The foundational rule: Evidence quality is determined by arguments that survive challenge, not by the letterhead they came on.

The Iraq WMD Test Case

In 2002, the expert consensus included the CIA, Colin Powell's UN presentation, bipartisan Congressional support, and major media amplification. The dissenters included weapons inspector Scott Ritter, some intelligence analysts, the Knight Ridder bureau, and ordinary citizens saying "this doesn't add up."

Under the traditional system, credentials won. Under the ISE framework, argument quality is what wins. The government's claim that aluminum tubes proved a nuclear program faced an immediate methodological challenge: the tubes were the wrong specification for centrifuges, and the Department of Energy's own experts said so. The claim that a single second-hand source confirmed mobile biological labs failed the replication test and the independent verification test at the same time. The dissenters' core argument, that "the evidence presented doesn't support the conclusion, and the details keep changing," was falsifiable, survived scrutiny, and its prediction (no WMDs found post-invasion) was validated.

In that framework the dissenters' arguments would have risen and the credential-backed claims would have sunk, with the whole record public. Because what mattered wasn't who had the fancy title. What mattered was whose arguments held up.

Related Pages

Truth Scores (overview) · Evidence (tiers and definitions) · Linkage Scores · Argument Scores from Sub-Argument Scores · ReasonRank · Media Truth Score · Media Genre and Style Scores · Importance Score · Objective Criteria Scores · Topic Overlap Scores

Contact me to contribute challenges, evidence, or corrections. | GitHub for the scoring implementation.

 

 

Comments (0)

You don't have permission to comment on this page.