edgar-sentiment
How it works

Methodology

This page explains the project in plain language first, then adds technical detail for readers who want the exact rules. This is not a forecast, not causation, and not trading advice.

At a glance

  • 9,697 scored filings in the published rollup
  • Main company board: 440 companies with at least 8 quarterly observations
  • 35 companies remain notable after the multiple-testing adjustment

Research question

Plain English

When management writes about the business in a quarterly filing, does the tone of that writing tend to move in the same direction as the company's earnings change for that same period?

Technical detail

Same-filing contemporaneous association between FinBERT-scored MD&A tone and YoY net income (primary) / revenue (secondary). Not prediction or causation.

Data

Plain English

The project covers current S&P 500 companies. Each company contributes recent annual (10-K) and quarterly (10-Q) SEC filings. The public rankings focus on quarterly filings with enough matched observations.

Technical detail

502 ticker rows; 499 unique SEC registrants after share-class consolidation. Primary board requires n≥8 on 10-Q net income pairs. Limited sample (n=6–7) appears on company pages only.

MD&A

Plain English

Management's Discussion and Analysis, or MD&A, is the section where management explains performance, trends, and risks in its own words. This project scores that section only, not the whole filing.

Technical detail

Extract Item 7 (10-K) or Item 2 (10-Q). Long MD&As are subsampled (evenly spaced, up to ~220 sentences).

Tone scoring (FinBERT)

Plain English

FinBERT is a language model trained for financial text. It estimates whether a sentence sounds positive, neutral, or negative. It is a structured reading aid, not a human judgment of emphasis or intent.

Technical detail

Model: ProsusAI/finbert. Sentence scores combine into a filing-level tone score.

Sentiment score

Plain English

Each filing gets one tone score from roughly −1 (more negative language) to +1 (more positive language). Higher means the MD&A language leaned more positive on average.

Technical detail

Filing score = mean of sentence scores, where each sentence score is positive share minus negative share.

Financial data

Plain English

For each filing we look up how net income (and, where comparable, revenue) changed versus the same quarter one year earlier. That year-over-year change is what we compare to tone. Before the correlations are calculated, each change is capped at the 1st and 99th percentile of all filings of that form for that metric. Every filing stays in the sample. The filing table still shows the uncapped percent. The scatter plots the capped percent so one extreme ratio does not flatten the other points.

Technical detail

Primary metric: 10-Q net income YoY. Secondary: 10-Q revenue (where comparable) and 10-K net income. Financials and Real Estate revenue comparisons are not used due to cross-company concept differences. The cap is the pooled 1st and 99th percentile, computed separately for 10-Q and 10-K, and for net income and revenue. Spearman, Pearson, agreement, sector filing-weighted and company-balanced results, and the multiple-testing adjustment all use the capped ratios. Spearman often stays the same, because a point that was already the most extreme usually keeps its rank. The cap does not fix a negative percent when the prior year was a loss. After the quarterly net-income cap, the chart axis can still run to roughly those percentiles.

XBRL

Plain English

XBRL is structured financial data reported to the SEC. This project uses those structured figures rather than scraping tables by hand.

Technical detail

SEC companyfacts XBRL for the same accession as the filing text.

Period matching

Plain English

Tone and numbers must come from the same filing period. Matching the wrong quarter can create a false relationship, so the project applies strict period rules.

Technical detail

Same accession; duration bands and report-date tolerance for period integrity. Combined / pooled 10-K+10-Q correlations are exploratory only and are not the public ranking metric.

Correlation overview

Plain English

A relationship (correlation) answers: when tone is more positive, do earnings changes also tend to be stronger? Two related measures are shown. Spearman is primary; Pearson is secondary, because extreme earnings swings can distort a simple straight-line fit.

Technical detail

Public primary: Spearman ρ on 10-Q NI. Secondary display: Pearson r with Fisher 95% CI when n is sufficient. See Spearman and Pearson.

Spearman

Plain English

Spearman measures whether more positive tone generally appears alongside stronger financial performance. It focuses on direction and is less affected by unusually large earnings changes. That is why Spearman is the primary public metric.

Technical detail

Rank-based association between MD&A tone and YoY net income for matched 10-Q observations.

Pearson

Plain English

Pearson measures the strength of a straight-line relationship. It can be more sensitive to unusually large earnings changes (for example, swings around near-zero prior income).

Technical detail

Linear Pearson r with Fisher 95% confidence interval when sample size allows. Shown as secondary detail on company pages.

Direction agreement

Plain English

Agreement counts how often tone and earnings simply moved the same way (both up or both down). Near-neutral cases are excluded so tiny noise does not count as a “move.”

Technical detail

Direction agreement after excluding near-neutral observations. Displayed as num / den (for example, 6 / 8).

Sample size

Plain English

Sample size is how many comparable quarterly filings enter the estimate. More observations usually make the estimate steadier, but a larger sample is not automatic proof of an important relationship.

Technical detail
  • n ≥ 10: more established sample (label only)
  • n = 8–9: usable; included on the default board
  • n = 6–7: limited; company page only
  • n < 6: insufficient for rankings

p-values

Plain English

The p-value measures how unusual a result this strong would be under a no-relationship assumption. Small p-values are a clue, not a verdict, especially when hundreds of companies are tested.

Technical detail

Two-sided p-values for Spearman and Pearson appear in collapsed statistical details on company pages.

FDR (multiple-testing adjustment)

Plain English

When hundreds of companies are tested, some can appear statistically notable by chance. False Discovery Rate, or FDR, adjusts for that problem.

440ranking-eligible companies tested
raw p < .05some look notable before adjustment
35remain after multiple-testing adjustment (q < .05)

FDR helps reduce the chance that the leaderboard is highlighting random statistical flukes. It does not prove the 35 survivors are economically important or causal.

Technical detail

Benjamini–Hochberg q among ranking-eligible companies. Badge when q < 0.05. Not proof or certainty.

Confidence intervals

Plain English

The confidence interval shows a range of plausible values for the estimated relationship. Wider intervals mean more uncertainty. It is not a guaranteed range.

Technical detail

Fisher 95% CI for Pearson r when n is sufficient.

Relationship labels

Plain English

On the site, Spearman values are also described in everyday language (for example, “strong positive”). These labels describe statistical relationship strength only. They do not measure company quality, management credibility, or investment value.

Technical detail
  • ρ ≥ 0.70: Strong positive
  • 0.40–0.69: Moderate positive
  • 0.20–0.39: Weak positive
  • −0.19–0.19: Little or no relationship
  • −0.39 to −0.20: Weak negative
  • −0.69 to −0.40: Moderate negative
  • ρ ≤ −0.70: Strong negative

Sector weighting

Plain English

Typical filing. That is the filing-weighted result: every filing counts.

Typical company. That is the company-balanced result: each company gets equal weight so large filers do not dominate.

These two views can differ (Industrials is a useful example). Neither alone is the full story.

Technical detail

Filing-weighted Spearman/Pearson and company-balanced Pearson on 10-Q NI. Winsorized Pearson shown when n≥20 for context.

How to read scatterplots

Plain English
  • Each dot represents one 10-Q filing
  • Left–right (x): MD&A tone (more negative ← → more positive)
  • Up–down (y): year-over-year financial change, after the pooled cap
  • Upper-right: more positive language with improving earnings
  • Lower-left: more negative language with worsening earnings
Technical detail

Primary company charts use 10-Q points only. The vertical position is the capped year-over-year ratio. Hover shows the uncapped percent. Sector charts pool 10-Q filings for the industry.

Limitations

Plain English
  • Near-zero prior earnings and loss↔profit flips can inflate Pearson
  • Capping extreme percent changes limits how far one point stretches the chart. It does not correct a misleading sign when the prior year was a loss.
  • MD&A extraction quality varies by filing layout
  • Model averages are not a human reading of emphasis or risk language
  • Sector filing-weighted vs company-balanced results can differ
  • Annual (10-K) history remains short for many companies
  • This site explores a same-filing relationship, not prediction

← Overview · About this project