Home > Page Design > Algorithms > Objective Criteria
Objective Criteria: Agreeing on the Yardstick Before the Fight
The Moving Goalposts Problem
You say: "The economy is great!"
I say: "The economy is terrible!"
Are we debating facts? No. You're looking at the stock market. I'm looking at inflation. We're not disagreeing about reality. We're disagreeing about which measurement tool counts as reality.
This is why most debates fail. Not because people can't agree on what's true, but because they can't agree on how to measure what's true. One person defines "good economy" as rising GDP. Another defines it as falling inequality. Another defines it as median wage growth. They're all measuring different things, calling it the same thing, and wondering why they can't convince each other.
It's like arguing about who's taller when one person uses feet and the other uses kilograms.
Before deep-diving into Pro/Con arguments, the ISE asks users to populate a table of shared metrics first. In a healthcare debate, for instance, users might agree that "cost per capita" and "life expectancy" are the primary rulers for measuring success. Those agreed-upon yardsticks then determine how the algorithm weights every subsequent argument. Change the yardstick and you change what counts as evidence. So we settle the yardstick first.
The golden rule of reason: Before asking "Who is right?", we must ask "How do we measure what is right?"
The Solution: Decide the Yardstick First
Every ISE topic page has a dedicated section where people propose and debate objective criteria. Think of it as a demilitarized zone where opponents agree on standards before they start fighting over conclusions.
The principle is simple: decide what counts as evidence before looking at the evidence. This is basic scientific method. You don't run an experiment, see what happens, then decide what you were measuring. You define your metrics first. Otherwise you're just cherry-picking whatever data supports your predetermined conclusion.
How It Works
Step 1: People Propose the Yardsticks
The community brainstorms specific, quantifiable metrics to evaluate the topic.
Topic: "Is the Economy Healthy?"
- GDP Growth Rate
- Median Real Wage Growth
- Labor Force Participation Rate
- Gini Coefficient (Inequality)
- Consumer Price Index
- Housing Affordability Index
Topic: "Is Candidate X Intelligent?"
- Standardized Test Scores (SAT/IQ)
- Vocabulary Complexity in Unscripted Speech
- Success Rate of Past Strategic Decisions
- Ability to Update Beliefs When Evidence Changes
Notice what this does. It forces specificity. You can't hide behind vague claims like "the economy is doing well" when six concrete measures might tell different stories. And once you've organized arguments this way for one president's economic record, it's trivially easy to apply the same structure to every president. Suddenly you can see how each side applies demanding standards to opponents while ignoring those same standards when evaluating their own.
Step 2: The Community Scores the Yardsticks
Not all measures are created equal. Some are better tools than others. People propose arguments about each criterion's quality across four dimensions:
1. Validity: Does this actually measure what we think it measures?
Example: Is the stock market a valid measure of human wellbeing?
- Against: Stock market mostly measures corporate profit. Can rise while median wages fall.
- For: Corporate health predicts future employment and economic stability.
Result: The engine weighs these arguments and derives a moderate validity score.
2. Reliability: Can different people measure this consistently?
- High reliability: Vocabulary grade level in speeches (objective, algorithmic)
- Low reliability: "Candidate seems smart" (subjective, varies by observer)
3. Independence: Is the data source neutral?
- Low independence: Study funded by an oil company on emissions safety
- High independence: NASA satellite data on atmospheric CO2
4. Linkage: How strongly does this metric correlate with the ultimate goal?
- High linkage: Median wage growth (directly affects most people's lives)
- Medium linkage: GDP growth (benefits spread unevenly)
- Low linkage: Stock market performance (affects investors primarily)
Goal: Determine how severe climate change is.
Average Global Temperature. Score: 85%
For: Direct physical evidence of heat retention
Against: Averages hide extreme regional variances
Glacier Mass Balance. Score: 92%
For: Ice melts only when heat is added over time (integrates data)
For: High reliability via satellite imagery
For: Hard to manipulate or misinterpret
Frequency of "Hot Days". Score: 60%
Against: Subject to recency bias and local weather patterns
Against: Low reliability in historical data comparison
Caveat: Affects humans directly but confounds weather with climate
Twitter Sentiment About Heat. Score: 15%
Against: Measures perception, not reality
Against: Heavily influenced by bots and viral events
Against: No correlation with actual temperature records
Illustrative scores, hand-picked to show the ordering the mechanism produces. Not engine output.
When the system calculates climate change severity, evidence based on glacier mass gets weighted heavily. Evidence based on Twitter sentiment gets filtered out almost entirely. This is exactly how reasoning should work. Better measures get more weight. Worse measures get less. The math is transparent and anyone can challenge it by proposing better criteria or showing why existing ones are flawed.
Why This Changes Everything
1. It Exposes Hidden Values
When two people argue about economic policy, they often think they're debating facts. They're not. They're debating values. One prioritizes GDP (growth). The other prioritizes the Gini Coefficient (equality). By forcing them to rank these criteria, we reveal that their disagreement is about what matters, not what's true.
That's actually progress. Values disagreements are negotiable. You can find compromises, balanced approaches, and Pareto improvements. Fake factual disagreements that are really values conflicts just generate endless frustration while disguising the real question.
The conflict pipeline reads the scored rows and classifies each dispute as factual, linkage, values, or mixed. "You two are actually arguing about values" is a computed readout on the belief page, not a moderator's opinion. In the seeded corpus, the UBI debate reads as a values conflict.
2. It Filters Noise Automatically
Once the community establishes that Twitter sentiment analysis is a low-quality metric for public opinion (due to bots, self-selection bias, and manipulation), any argument relying on Twitter data automatically gets downgraded. You don't have to re-litigate the reliability of Twitter polls every time someone cites one. The criteria score is already calculated. The algorithm already knows it's weak evidence. The math handles it permanently.
This is core negotiation theory (Fisher and Ury). If you agree on criteria first ("we will choose the vendor with lowest cost per unit"), the final decision becomes calculation rather than ego battle. Agreeing on measurement standards before looking at data removes most opportunities for motivated reasoning. You can't move the goalposts if you committed to them publicly before seeing which direction they pointed. The conflict resolution pipeline runs on exactly this foundation: shared criteria in, shared interests, the primary conflict pair, and compromise candidates out.
How Criteria Get Scored (The Recursive Part)
Criteria scores are themselves determined by arguments, the same recursion the engine runs everywhere else.
Question: "Is glacier mass balance a good measure of climate change?"
Arguments pushing the score HIGH (80-95%):
- Ice responds to sustained heat, not daily fluctuations (integrates signal over time)
- Satellite measurements are objective and replicable
- Glacier retreat correlates strongly with atmospheric CO2 (causal linkage established)
- Independent verification possible across different glacier systems worldwide
Arguments pushing the score MEDIUM (50-70%):
- Some glaciers affected by local precipitation changes unrelated to global warming
- Historical baseline data less precise than modern satellite measurements
- Doesn't directly measure atmospheric temperature (proxy measurement)
Arguments pushing the score LOW (below 50%):
- Glaciers are only one component of Earth's ice systems
- Regional variations make global aggregation challenging
Each argument gets evaluated for evidence quality, logical connection, and importance. The preponderance of well-supported arguments determines the final criteria score. In this case: ~92%. Glacier mass balance is an excellent measure of sustained warming. (Illustrative worked example; the number is hand-computed to show the balance, not engine output.)
Integration with the Broader System
Objective Criteria are the scales upon which all other evidence gets weighed.
Truth Scores get calculated by comparing claims against high-scoring criteria. A claim measured by excellent criteria (glacier mass, peer-reviewed studies) gets higher confidence than a claim measured by poor criteria (anecdotes, Twitter polls).
Importance Scores are anchored here. Every importance sub-debate ties back to the objective criteria debate for its belief, and the anchor criteria are seeded into each one as falsifiability tests: the evidence that would move the multiplier, named in advance. Nobody gets to say a point matters without also saying what would prove it doesn't.
Cost-Benefit Analysis requires agreed-upon units of measurement. Are we measuring in dollars? Quality-adjusted life years? Lives saved? You can't do meaningful cost-benefit analysis without establishing criteria first.
Assumptions are made explicit by the criteria layer. Selecting GDP as your primary economic measure assumes growth is the goal. Selecting the Gini Coefficient assumes equality is the goal. Making those assumptions visible prevents hidden disagreements from quietly sabotaging analysis that looks objective on the surface.
Common Objections
"But some things can't be measured objectively!"
True. That's exactly why we score criteria quality. Subjective measures get low scores. Arguments relying on them get downweighted accordingly. If you're debating something genuinely unmeasurable, the criteria layer makes that explicit. "We have no good way to measure this" is itself valuable information. It tells you to hold your conclusions lightly.
"Who decides which criteria matter?"
The community, through evidence-based argument. Someone proposes a criterion. Others propose arguments for or against its validity, reliability, independence, and linkage. The ReasonRank algorithm synthesizes the competing arguments into a score. No individual decides. The collective evaluation, weighted by argument quality, decides.
"Can't people game this by proposing biased criteria?"
They can try. But they have to defend their criteria publicly against informed opposition using verifiable evidence. "Stock market performance" as a measure of human welfare gets challenged with evidence that it's weakly linked to median quality of life. The bias becomes visible in the scoring. Gaming a transparent system is dramatically harder than gaming one with no criteria evaluation at all.
What This Looks Like in Practice
Without Objective Criteria: Endless debates where people talk past each other. "The economy is great!" "No it's terrible!" "Stock market is up!" "Wages are flat!" Nobody wins because nobody agreed what winning means.
With Objective Criteria: The topic page for "Economic Health" lists six criteria with scores. Users can filter by priority: growth-focused view, equality-focused view, or median-impact view. The data doesn't change. The framing adapts to what you care about. The disagreements become transparent.
You see: "By GDP and stock market measures (weighted for growth priority), the economy scores 85%. By wage growth and inequality measures (weighted for equality priority), it scores 45%."
Now you're having an honest conversation. The facts are clear. The values are explicit. The remaining disagreement is about what we should optimize for, not about what reality is. That's a conversation worth having.
Why This Matters for Democracy
Most political dysfunction comes from fake disagreements. People argue about "facts" when they actually disagree about values. Or they argue about values when they actually disagree about what data is reliable. Objective Criteria separate these layers:
- Criteria layer: What counts as good measurement?
- Data layer: What do those measurements actually show?
- Values layer: Which measurements should we prioritize?
Once you separate them, most debates get dramatically simpler. Some questions are empirical (what does the data show given these criteria?). Some are normative (which criteria should we care about most?). Mixing them creates endless confusion. Separating them creates the possibility of resolution. The conflict pipeline computes which layer a given dispute is actually stuck in, so the diagnosis is on the page instead of in the eye of the beholder.
Contribute
The ISE is infrastructure, not a publication. The structure exists. The scoring logic runs. What fills it in is community knowledge: argued, evaluated, and scored the same way every other belief on the platform gets evaluated.
If you see a major contested topic where no one has agreed on measurement standards, that's the most valuable place to start. Propose criteria. Argue for their validity. Challenge weak ones. The platform does the math. You provide the reasoning.
Get involved, or go straight to the code. The most important debates in the world are stalled not because we lack data, but because we haven't agreed on what the data means. That's a solvable problem.
Examples: Setting the Yardstick Before the Fight
The fastest way to end a pointless argument is to define what "good" means before looking at the data. Here is how the Idea Stock Exchange turns contested questions into measurable criteria, applied consistently, regardless of who benefits.
1. Politicians (Senate, Congress, President)
Ignore the soundbites. Measure what actually predicts effective governance.
- Independence from concentrated money: Percentage of campaign funding from small-dollar individual donors versus PACs, corporate bundles, and dark money. Higher independence = higher score.
- Legislative effectiveness: Ratio of bills sponsored that advanced out of committee versus total bills introduced. Prolific introducers who accomplish nothing score low.
- Bipartisan track record: Percentage of passed legislation that required cross-aisle co-sponsorship. Measures actual deal-making rather than claimed willingness to compromise.
- Accuracy of public claims: Independent fact-check ratings aggregated over term. Persistent misrepresentation of data lowers the score.
- Constituent responsiveness: Town halls held, response rate to constituent inquiries, accessibility outside of election years.
- Credible findings of misconduct: Adjudicated ethics violations, confirmed criminal conduct, or sustained formal investigations. Applied equally to all Candidates.
Additional Criteria for Executive Roles (President, Governor)
- Organizational track record: Have they successfully managed large institutions without driving them into insolvency, regulatory intervention, or mass litigation?
- Crisis decision quality: Retrospective evaluation of major decisions made under uncertainty. Did the reasoning hold up? Were costs acknowledged upfront?
- Appointee quality: Track record of selecting qualified, confirmed, non-scandal-plagued subordinates. Executives are responsible for who they choose.
2. Everyday Products
- Total cost of ownership per mile: Purchase price amortized over expected lifespan, plus average fuel and maintenance costs. Sticker price alone is misleading.
- Environmental cost: Lifetime carbon footprint including manufacturing, fuel source, and disposal.
- Safety per mile driven: NHTSA and IIHS crash ratings weighted by miles driven in that vehicle class.
- Reliability: Consumer Reports long-term reliability data, not manufacturer claims.
Cell Phones
- Privacy: Independent audit scores for data collection practices. How much personal data is collected, shared with third parties, or sold to advertisers?
- Repairability: iFixit repairability score. A phone you can't fix is a phone designed to be replaced.
- Software support lifespan: How many years of security updates are guaranteed? A cheap phone that stops receiving patches is a security liability.
- True cost over four years: Device cost plus plan cost plus accessory lock-in, not just the monthly payment.
3. Career and Life Paths
- Automation displacement risk: Oxford/McKinsey automation probability index for the specific role, updated to reflect current AI capabilities, not the general industry.
- Salary adjusted for cost of living and hours worked: $120,000 in Manhattan working 70 hours a week is a worse deal than $75,000 in Denver working 45.
- Reported burnout rate: Gallup and sector-specific survey data on sustained job satisfaction, not just starting enthusiasm.
- Mobility and optionality: Does this path open doors or close them? Skills transferability across industries.
- Long-term income trajectory: Median earnings at 10 and 20 years, not just entry-level salary.
4. Lifestyle Choices
Marriage and Long-Term Partnership
- Long-term life satisfaction: Longitudinal survey data comparing sustained partnership versus alternatives at ages 50, 65, and 80, not snapshot happiness measures.
- Health outcomes: Mortality rates, cardiovascular health, and mental health outcomes across relationship structures, controlling for socioeconomic status.
- Financial stability: Comparative wealth accumulation and bankruptcy rates over a 30-year period.
Having Children
- Life satisfaction at 65+: Longitudinal data on fulfillment, loneliness, and regret among parents versus non-parents past retirement age, the period when the actual long-term costs and benefits are clearest.
- Mental health trajectory: Depression and anxiety rates across parenting stages, disaggregated by support structures and socioeconomic conditions.
- Financial impact: Lifetime wealth differential, including opportunity costs of time and career interruption.
Related Scores
The neighboring scores in the pipeline each have their own page: Argument scores from sub-argument scores, Evidence, Truth, Importance, Linkage, Linkage Score Code, Topic overlap, Media Truth, Media Genre and Style, and Book Logical Validity. The full index is on the Scores page.
Comments (0)
You don't have permission to comment on this page.