The Super Data Collectors: How Similarweb Records the Browsing Behavior of the Entire WorldArtificially generated
Artificially generatedArtificially generatedArticle 50 of the AI Act requires artificially generated image, audio and video content to be marked as such, detectably and machine-readably. The symbol is the European Commission's official base icon. Using it is voluntary, the marking duty is not.Machine-generated: Text · Image. With human direction.McGrinsey Living Intelligence SystemsMcGrinsey Living Intelligence SystemsLI: Living intelligence of every kind. McGrinsey combines conventional intelligence and novel intelligence into living intelligence systems. These symbiotic systems of machine and human intelligence work together to supply the finest distillates of insight, for every intelligence.
MCG Research · Data 01 · August 2026

The Super Data Collectors

How Similarweb records the browsing behavior of the entire world.

An Israeli startup provides visitor numbers for one hundred million websites without operating a counter on a single one of them. This is not magic and not a scandal, but a very cleverly built machine. This report takes it apart: which four sources feed into it, how three million browsers become a statement about one hundred million websites, where the panel comes from, who built the company, and how accurate the result is in the end.

01 · The Question

How does someone know how many visitors a third-party website has?

Every second pitch deck contains a Similarweb number. Competitive analyses, due diligence, media planning, hedge fund models: somewhere in there sits a bar chart with monthly visits to a website to which the author has no access whatsoever. And it works surprisingly well.

The question behind it is a nice technical question, and it has a clean answer. The number is not a measurement, it is an extrapolation. Just like an election forecast: a sample is observed, corrected, calibrated against known truths, and then stretched to the total population.

Anyone who understands an election forecast understands Similarweb. It's the same four work steps, the same sources of error, the same kind of caution when reading the result. The only difference is the sample: instead of a thousand people called, it's millions of browsers.

This report goes through the machine from the beginning. First the four inputs, as the company itself describes them. Then the calculation method. Then the origin of the panel, which is its own and quite interesting story. After that the company, the numbers, and the accuracy.

First the scale: Similarweb is not a special case in this industry, but the best-documented example. What is written here applies in principle to every provider of market data from behavioral measurement.

100million+
Websites in inventory
This many websites are covered by Similarweb according to its own statements, plus 4 million apps and 235 million product articles.
10billion/day
Digital signals
Daily data inflow according to own statement. Around 2 terabytes are analyzed per day.
3.32million
Browsers in store
Sum of users of all four Chrome extensions listed under Similarweb accounts in the store. Counted ourselves on August 7, 2026.
250million $
What it's worth
Data points in money: around 283 million dollars annual revenue in 2025 from selling these analyses.
02 · The Four Inputs

What flows in, and what each source is good for

The middle column of the following table is not a summary, but the translation of the wording from the manufacturer's page Our Data. The right column explains the role in the calculation process. It is important that the four sources do not do the same thing: one delivers truth, one delivers behavior, one delivers structure, one fills gaps.

Fig. 01 The four data sources

Order as on the manufacturer's page. How strongly each source contributes to a single number is not published.

SourceOwn descriptionRole in calculation process
Website and app operators Directly measured data from their own analytics software (e.g. Google Analytics) from millions of websites and apps. The calibration standard. Operators voluntarily provide their real numbers in exchange for comparative values. Only here does Similarweb know the truth, and only against this can the model be adjusted.
Contributors Network Anonymous traffic data collected from Similarweb products installed on millions of devices worldwide. The panel. Browser extensions and apps that record which addresses are accessed. The actual raw material: only here does behavior emerge across third-party websites.
Public data Publicly available data (e.g. Wikipedia, census data), algorithmically captured and indexed. The map. Crawlers provide structure, categories, search result pages, app store positions, and population numbers. Says little about visitors, but a lot about what to compare them against.
Partnerships Pre-analyzed data from global partners such as DSPs, internet providers, measurement service providers, and credit agencies. The gap filler. Purchased movement data from sources that the own panel does not reach: other age groups, other countries, other devices. Greatly expanded after 2018.
Source: similarweb.com/corp/ourdata, accessed on August 7, 2026. Translation McGrinsey. Right column: own classification.
The trick is redundancy. Because multiple sources see the same phenomenon from different sides, they balance each other out: if the panel has a gap in older users, an internet provider fills it. This is methodologically strong and the reason why the system has survived changes at Google, Apple, or in legislation. However, it has a price for the reader: as long as the weights are not published, it is not possible to trace where a single number actually comes from. You have to believe the model or test it against your own data.
03 · The Calculation Method

How three million browsers become one hundred million websites

This is the core of the whole thing, and it is surprisingly comprehensible. Four steps, each with its own error, which the next step partially captures again.

4
Work steps from observation to published number. The same four as in every election forecast.
Step 01
Collect
What does the panel access?
Each device in the network reports the accessed addresses. This creates a raw count: this many panel devices were on this domain in July.
Step 02
Correct
Who does the panel see too often?
A panel of browser extensions is not a random sample. It overrepresents technically savvy people, certain countries, desktop over mobile. Each group therefore gets a weight, calculated against population and market data.
Step 03
Calibrate
Where do we know the truth?
The decisive step. Real analytics numbers are available for tens of thousands of websites. The model calculates the estimate for these sites, compares it with the truth, and corrects itself until it fits.
Step 04
Extrapolate
And all the rest?
The calibrated function is applied to all other websites for which no one knows the truth. This is exactly where the number you see in the tool is created.
Step 03 explains almost all characteristics of the result. Calibration works best where many calibration points exist. Large websites are more likely to provide their analytics than small ones, so the model is better calibrated for large websites. This is exactly what the accuracy studies in section 09 also measure: the error grows the smaller the site becomes. This is not sloppy work, this is the inevitable characteristic of a calibrated model.

The second peculiarity follows from step 02. A panel consisting of browser extensions sees the world through a desktop browser. Everything that bypasses it, it sees poorly: apps, embedded views in social networks, and increasingly answers provided by an AI system without anyone accessing a website. That's why the company has been buying partner data and entire companies for years. The panel is not the product, it is the calibration basis that must be constantly updated.

04 · Where the panel comes from

Millions of people contribute measurements without that being their goal

Now for the most difficult ingredient. A panel of this size doesn't arise from people voluntarily signing up for market research. It arises from a product being useful and recording on the side.

This isn't a Similarweb invention, it's the basic model of behavioral measurement since TV ratings have existed. What's new is the scale: instead of a few thousand households, it's millions of devices, and the compensation isn't cash but a free tool.

The origin of this panel has been documented over ten years by security researchers, universities, and an ad blocker manufacturer. The following timeline is the sober summary of these works.

42
Chrome extensions with a combined 8 million installations in which the same recording library was found in 2016. None of them were called Similarweb.
Fig. 02 Ten years of documented findings

Only incidents with published technical analysis. The reach figures cannot be added, they partially overlap.

DateFindingReachWho
03/2016 42 Chrome extensions contain the library upalytics and send every visited address, every search query, and even addresses from internal corporate networks to nine domains, all registered through an anonymization service. Among them similarsites.com. The same library is embedded in the Similarweb extension itself. 8.0M Michael Weissbacher, Northeastern University
01/2017 Similarweb acquires the design extension Stylish, with which users change the appearance of third-party websites. Starting with the same version, the Chrome version begins recording. ~2.0M Acquisition, public
07/2018 Stylish sends every complete address along with a permanent identifier to api.userstyles.org, double base64-encoded. Complete address means in practice also: login tokens, password reset links, search results. Google and Mozilla remove the extension from their offerings within two days, a revised version returns in August. ~2.0M Robert Heaton
07/2018 Nine tools from a newly founded Delaware company Big Star Labs send every visited page to their own servers. AdGuard determines the data format resembles that of Stylish, but explicitly calls a connection unconfirmed. 11.0M AdGuard
12/2025 The Similarweb extension reads conversations from ChatGPT, Claude, Gemini, and Perplexity, via a dynamically loaded configuration file with its own evaluation logic for each provider. The capability was added with an update in May 2025. The researcher calls the pattern Prompt Poaching. On January 1, 2026, Similarweb adds an explicit consent dialog for this. 1.0M John Tuckner, Secure Annex
02/2026 Stylish transmits addresses and AI conversations through five-layer obfuscation: URL encoding, double base64, column swapping, AES-256-CBC with key embedded in source code, finally base64 again. ~2.0M James Arnott
05/2026 Systematic scan of 287 extensions with a combined 37.4M users, roughly one percent of all Chrome users worldwide. Similarweb appears as the most frequent recipient via multiple pathways. The report also attributes Big Star Labs (3.7M users) to the company. 37.4M Q Continuum
The 37.4M from the 2026 scan refers to all examined extensions, not only those with Similarweb connections. The attribution of Big Star Labs comes from this report and is not confirmed by Similarweb.
The pattern has been stable for ten years, and it has an economic logic. A tool that advertises it will collect your browsing behavior is installed by hardly anyone. A tool that changes designs or blocks pop-ups is installed by millions. The utility is real, the recording runs alongside. This coupling is precisely what creates a panel of the size needed for steps 02 and 03. Smaller panels produce worse numbers.

From 2018 onwards, the situation changes. Google and Mozilla tighten their rules, the European General Data Protection Regulation takes effect, and Similarweb audibly restructures: more purchased partner data, less dependence on the Store. In the 2021 agreement with security company Check Point, it even explicitly states that Stylish data is not part of the exchange. In its securities filings, the company still lists dependence on the Contributory Network as its own risk factor: platform rules change frequently and are enforced inconsistently.

05 · Counted up

How many extensions does Similarweb operate today?

The obvious answer is two: the traffic extension and the sales extension, both bearing the name. Instead, on August 7, 2026, we extracted the publisher IDs from the Chrome Web Store, meaning we compared accounts, not product names. This adds two more entries.

Stylish still appears under the publisher account similarweb-ltd and with two million users is the largest item. Similar Sites runs under similargroup, the old company name. Neither carries the Similarweb name in the Store listing.

Fig. 03 Chrome extensions under Similarweb accounts

User numbers as shown by the Chrome Web Store, retrieved on August 7, 2026. The Store rounds to even increments.

Stylish: Custom themes
2,000,000
Similarweb: Traffic & SEO Checker
1,000,000
Similar Sites
300,000
Similarweb Sales Extension
20,000
ExtensionPublisher in StoreBears the name?Users
Stylish: Custom themes for any websitesimilarweb-ltdno2,000,000
Similarweb: Website Traffic, AI Traffic & SEO CheckerSimilarwebyes1,000,000
Similar Sites: Discover Related Websitessimilargroupno300,000
Similarweb Sales ExtensionSimilarwebyes20,000
Chrome total3,320,000
In addition there are versions for Edge, Firefox and Opera that the manufacturer itself mentions, as well as mobile apps and the purchased partner sources. The Store only shows Chrome numbers.
This count captures only the visible part. The 42 extensions from 2016 bore third-party names and were found because someone compared the library in the code, not the imprint. From the outside, without such a scan, it's impossible to say how many devices in total feed into the network. Similarweb itself only mentions "millions of devices," without a number.
06 · What exactly is collected

The answer is in the terms, and it is surprisingly open

For the question of what data actually flows, you don't need a security researcher. You just have to read the right file, and that's the small difficulty: Similarweb's privacy policy explicitly does not apply to the browser extension. It says so in its own sentence.

The declaration that applies to the extension is on a different domain: similarsites.com. There it is stated very clearly who is responsible and what is collected. This is not a leak and not a discovery, this is the text that is agreed to upon installation.

The last line in the table is the most economically important: the result may be passed on to business customers. This is exactly the product that Similarweb sells.

Fig. 04 What the extension collects according to its own declaration

Verbatim from the privacy policy at similarsites.com, accessed on August 7, 2026.

CategoryOwn wording
Responsible"Similarweb Ltd., incorporated under the laws of the State of Israel, is the controller"
Addresses"URLs accessed or visited, pages on which advertising was seen or clicked, advertising URL"
History"Clickstream data, search history and information about interactions"
Search"Data from search results pages (search term, order or rank of results)"
AI inputs"Prompts, queries, content and other inputs that you enter into or submit to certain artificial intelligence tools"
Sharing"We may 'sell' or 'share' these categories to business customers so that our business customers can better understand consumer behavior"
Translation McGrinsey. The quotation marks around sell and share are in the original, they refer to the legal definition in California consumer protection law.
The calculation method can be read back from the list. Complete addresses are needed to distinguish subpages and entry points. Search history and results page rank are needed to demonstrate search engine visibility. AI inputs are needed for the company's latest product, the analysis of traffic from chatbots. Each line in this table corresponds to a column in the sold report.
07 · The company

From jewelry store to stock exchange: Or Offer

Or Offer, born in 1983, grows up in Tzur Hadassah near Jerusalem, as a computer kid with parents who design jewelry and struggle with the uncertainty of this profession. After school, intelligence unit 8200 wants him as a programmer. He declines and goes to Oketz, the canine unit, among other things to a post in the Gaza Strip. About this time he later says it taught him to stand up to experienced people as a young person.

Afterwards he studies business administration part-time at IDC Herzliya and builds two jewelry stores, in Zichron Ya'akov and Netanya, plus an online shop. The chain is called Maya Offer, after his mother. His parents push him towards high-tech.

The trigger is a name: a supplier mentions the designer David Yurman, who combines stones with silver. Offer wants to find comparable products online and realizes that there is no tool for this. Similarweb is created in 2007 as a browser extension that suggests similar websites to a website. Hence the name, and hence also the panel: data collection is originally the byproduct of a search function.

He is 23 and works for over a year without salary. In 2007 an important collaborator leaves the project because he doesn't believe in the product. At 24 Offer thinks about quitting and stays: "I had nothing else to do, so I continued."

In 2009 Yossi Vardi gives him one million dollars, for about five percent. The company lives on this for five years, with seven or eight people. In 2011 comes the pivot that decides everything: away from the consumer product, towards an analysis platform for companies and investors. The consumer product doesn't disappear in the process, it changes roles. From now on it is no longer the product, it is the measuring point.

Offer was sole founder and considers this an advantage: "When there are multiple founders, each needs a role to feel important. Precisely because I was alone, that saved me a lot of politics."

In 2014 Naspers leads a round of 18 million dollars. In May 2021 the company goes public on the New York Stock Exchange, valued at 1.6 billion dollars, and raises around 180 million dollars. Headquarters is still Givatayim near Tel Aviv, plus New York and London.

Fig. 05 Revenue 2019 to 2025, in million US dollars

The growth rates of over 40 percent occurred in the years around the IPO. Since 2023 growth has been stable in the low teens.

2019
70.6
2020
93.5
2021 · IPO
137.7
2022
193.2
2023
218.0
2024
249.9
2025
282.6
YearRevenue (million $)Growth
201970.6
202093.5+32.4 %
2021137.7+47.3 %
2022193.2+40.4 %
2023218.0+12.8 %
2024249.9+14.6 %
2025282.6+13.1 %
Fiscal years. Source: Annual reports and quarterly announcements of the company.

Breadth was acquired

The company's own panel measures web traffic in the browser. Everything beyond that, app data, advertising data, search engine rankings, came through acquisitions. Eight documented purchases in ten years, purchase prices almost never published. The list reads like a map of the blind spots that a browser panel has.

Fig. 06 Acquisitions
YearCompanyWhich gap was closed with it
2014TapDogEarly-stage startup, shares and cash
2015SwayyContent discovery
2015QuettraMobile analysis, i.e. the first blind spot of the browser panel
2017StylishAround two million browsers at once, i.e. pure panel size
2021Embee MobileMobile panel data
2022Rank RangerSearch engine rankings and interfaces
2024AdmetricksAdvertising data
202442mattersApp data
2025The Search MonitorPaid search, brand protection, affiliate program control
Stylish is not in the company's official acquisition list. The purchase is documented through the 2017 announcement and through the publisher account in the Chrome Web Store.
08 · Where the company stands today

The raw material becomes training material

On May 13, 2026, the board of directors opened the search for a new CEO. Or Offer is leaving by mid-2027, after almost exactly twenty years. "Similarweb was my life's work," he says, it is "the right moment." Board chairman Harel Beit-On calls the result "a unique data asset."

The market capitalization in May 2026 was around 270 million dollars, a good 56 percent below the beginning of the year and far below the 1.6 billion of the IPO. Revenue is growing, the stock price is not. That is the framework in which to read the next step.

In the first quarter of 2026, Similarweb signed a seven-figure contract for training data for a large language model, with an existing major customer. A second is announced. Together with other multi-year deals, the company cites 47 million dollars in total contract value.

This creates a remarkable loop. The panel records how people talk to AI systems. The analysis of this is sold to the manufacturers of such systems. And a third product sells brands the observation of how these systems recommend their products. The same raw material, three markets.

For the reader, the interesting thing about this is the direction: behavioral data is currently in the process of becoming a raw material for models rather than an analysis tool. Similarweb is an early, well-documented case of this.

282.6M $
Revenue 2025
Plus 13 percent. First quarter 2026: 73.9 million dollars.
6,128
Customers
As of end of 2025. Of these, 454 with at least 100,000 dollars in annual revenue.
~1,120
Employees
End of 2025, slightly down from 2024.
−56%
Market cap since beginning of year
As of May 2026, around 270 million dollars market capitalization.
09 · How accurate is the result?

The error is large, and it doesn't always point in the same direction

There are several independent comparisons against real Analytics access. They contradict each other in sign, and that is precisely the result. Anyone expecting a single error figure has misunderstood the matter: the error depends on how large the site is, what type of site it is, and how well Analytics itself measures on that site.

Fig. 07 Independent accuracy comparisons

All against Google Analytics as reference. Different samples, different years, different signs.

StudySampleFinding
Omniconvert, 20261,787 shopsSessions around 94 percent too high. Error systematic, not random. Large sites significantly more accurate than small ones.
PLOS One, 202286 websitesVisits 19.4 percent too low, unique visitors 38.7 percent too low, bounce rate 25.2 percent too high.
SparkToroProvider comparisonFor sites with 5,000+ users most frequently within ±30 percent and thus better than all competitors. Below 5,000 users worse.
The signs contradict each other because the samples are different: shops behave differently than editorial sites, and Analytics itself measures too little depending on consent rates and ad blockers, so it is not a perfect benchmark.
The useful reading is the relative one. Whether a website has 400,000 or 700,000 visits, Similarweb doesn't tell you reliably. Whether it is two orders of magnitude above yours, whether its trend is rising against yours, which channels feed it and from which countries it comes, for that it works very well. The reason is in section 03: the calculation method optimizes for comparability across many sites, not for the exact number of a single one.
10 · Assessment

Four sentences to get the picture right

First: nothing is being accused here. There is no case in which a legal violation has been proven to an authority. This report describes a mechanism, it does not accuse.

Second: Similarweb contradicts the 2016 presentation. The company has stated that these are "repeated foreign reports from 2016 that were already baseless at the time and were refuted many years ago." This sentence stands here as it was made.

Third: the company has visibly restructured. After 2018, away from dependence on the Store, toward purchased partner data, with explicit exclusion of Stylish data in the Check Point agreement and a consent dialog for AI content since January 2026.

Fourth: this is an industry model, not an isolated case. Behavioral panels from consumer products have existed since television ratings. The 2026 scan found 287 extensions, not four. Similarweb is the best-documented example here because it is publicly traded and therefore must be accountable.

And the point at which the whole thing dissolves: there is no method to estimate third-party traffic without observing someone somewhere. The alternative to a panel is not a better panel, but no number at all. Anyone who uses the number uses the method with it.

An extrapolation is only as good as what it was calibrated against.
11 · What this means in practice

For your analyses

Use Similarweb for ratios and trends, not for absolute numbers in a calculation on which something depends. The documented error ranges from half to double and reverses sign depending on site type.

For small sites, be doubly cautious

Below about 5,000 monthly users it becomes unreliable, for a structural reason: there are too few calibration points and too few panel hits there.

For your own site

Your numbers are already in the panel, and your competition can read them. This is not an attack, but the normal case. Plan with it.

For your work devices

A browser extension sees everything that happens in the tab. On a computer with customer data, contracts, or AI conversations in the browser, every installed extension is a decision about confidentiality. Counting them once and checking the publisher for each one is a matter of ten minutes.

Voices · Interviews and original statements
  • Or Offer, extensive portrait interview (Calcalist/CTech). Childhood, Oketz, the jewelry chain, the David Yurman moment, the years without salary, Vardi's million, the pivot in 2011. The source for almost all quotes in section 07. calcalistech.com
  • Or Offer in conversation with Forbes, February 2017. Entrepreneurship, crises, why a pivot is "like a divorce." forbes.com
  • Or Offer on his departure, May 2026 (CTech). "Similarweb was my life's work," plus the current business figures and the stock price decline. calcalistech.com
  • Original announcement on succession, May 13, 2026. Wording from Or Offer and board chairman Harel Beit-On. ir.similarweb.com
  • Marta Sulkiewicz, VP Emerging Solutions, on the AI business (MediaPost, May 2026). On the training data contract and what brands are asking today: "Every day we are asked if we can see purchase completions in AI agents and bots." mediapost.com
  • Michael Weissbacher, Northeastern University, on the 2016 investigation. The initial description of the panel from a security research perspective. mweissbacher.com
  • John Tuckner, Secure Annex, on Prompt Poaching. "Prompt Poaching has arrived to siphon off your most sensitive conversations, and browser extensions are the way there." secureannex.com
  • Quarterly reports with Or Offer in original audio. All transcripts and recordings of the conference calls. ir.similarweb.com
Sources

Everything this report is based on

Fourteen sources, each with what it contributes and the full address. They are intentionally set large and expanded: those who don't need them scroll past, those who want to verify shouldn't have to search.

  1. Similarweb: Our Data
    The four data sources in the manufacturer's own words, plus the size specifications: over 100 million websites, over 4 million apps, 235 million product articles, 10 billion signals and 2 TB per day, 200 data scientists and 50 PhDs.
    similarweb.com/corp/ourdata/
  2. Similarweb: Privacy Policy
    Contains the crucial sentence that this policy explicitly does not apply to the app and browser extension. This begins the trail to the policy that actually applies.
    similarweb.com/corp/legal/privacy-policy/
  3. SimilarSites: Privacy Policy
    The policy that actually applies to the extension. Names Similarweb Ltd. as the controller and lists complete URLs, clickstream, search history and AI inputs, plus the sale to business customers. The basis for Fig. 04.
    similarsites.com/privacy-policy
  4. Similarweb: Form 20-F, Fiscal Year 2021 (SEC)
    The risk disclosure in which the company itself identifies the Contributory Network's dependence on extensions and apps in third-party stores as a business risk. The US Securities and Exchange Commission server rejects automated requests, but the document is normally accessible in the browser.
    sec.gov/Archives/edgar/data/1842731/000184273122000013/smwb-20211231.htm
  5. Similarweb Investor Relations: Quarterly Results
    Revenue 2025, customer count, customers with annual revenue of $100,000 or more, the first quarter of 2026, and the seven-figure contract for training data.
    ir.similarweb.com/financials/quarterly-results
  6. Michael Weissbacher, Northeastern University: These Chrome extensions spy on 8 million users (March 2016)
    The initial description. 42 extensions with the upalytics library, 8 million installations combined, nine recipient domains registered through an anonymization service. The starting point of the entire chain of evidence.
    mweissbacher.com/2016/03/31/these-chrome-extensions-spy-on-8-million-users/
  7. Ex-Ray: Detection of History-Leaking Browser Extensions (ACSAC 2017)
    The scientific continuation from Northeastern University and University College London. 10,691 extensions examined, 212 identified as history-leaking, solely from the pattern of network traffic.
    mweissbacher.com/publications/acsac_exray.pdf
  8. Robert Heaton: Stylish steals all your internet history (July 2, 2018)
    The technical analysis that led to removal from the stores. Complete addresses including persistent identifier to api.userstyles.org, double base64-wrapped, linkable to the account via the session cookie.
    robertheaton.com/2018/07/02/stylish-browser-extension-steals-your-internet-history/
  9. AdGuard: Big Star Labs spyware campaign affects over 11.000.000 people (July 2018)
    Nine tools from a Delaware company, recipient domains, privacy policies as image files. Contains the explicit caveat that the suspected connection to Similarweb is not confirmed. That's exactly why it's also listed here as unconfirmed.
    adguard.com/en/blog/big-star-labs-spyware.html
  10. John Tuckner, Secure Annex: Prompt poaching runs rampant in extensions (December 2025)
    The original analysis on intercepting AI conversations: dynamically loaded configuration file with custom evaluation logic per provider, for ChatGPT, Claude, Gemini and Perplexity. Coins the term Prompt Poaching.
    secureannex.com/blog/prompt-poaching/
  11. Q Continuum: Report on spying browser extensions (2026)
    The largest systematic scan to date: 287 extensions with 37.4 million users, 930 CPU-days of measurement time, methods and individual findings disclosed. Also attributes Big Star Labs to the company, which Similarweb has not confirmed.
    github.com/qcontinuum1/spying-extensions
  12. Consumer Rights Wiki: Entry SimilarWeb
    The continuous timeline of incidents through May 2026, including James Arnott's analysis of the five-level obfuscation and the classification of both extensions as confirmed exfiltrators.
    consumerrights.wiki/w/SimilarWeb
  13. Globes: Similarweb's controversial route to Wall Street (2021)
    The most comprehensive journalistic investigation. Contains the company's statement verbatim and the agreement with Check Point, in which Stylish data is explicitly excluded.
    en.globes.co.il/en/article-similarwebs-controversial-route-to-wall-street-1001376912
  14. Accuracy studies: Omniconvert (2026) and PLOS One (2022)
    The two comparisons against real analytics access from section 09. Omniconvert across 1,787 shops, PLOS One across 86 websites, with opposite signs. Both are linked because only the contradiction carries the statement.
    omniconvert.com/blog/we-analyzed-1787-ecommerce-websites-similarweb-google-analytics-thats-we-learned/ journals.plos.org/plosone/article?id=10.1371/journal.pone.0268212

McGrinsey's own work: The user numbers and publisher accounts of the four Chrome extensions in Fig. 03 were read directly from the store entries on August 7, 2026. The Chrome Web Store rounds user numbers to smooth levels, so the sum of 3,320,000 is correspondingly rough. The calculation method in section 03 is a reconstruction from the published methodology and the academic literature on panel extrapolation; the company does not publish it in detail. All translations are ours. Where sources contradict each other, the contradiction is stated in the text.

Connection
Data 01 is the beginning of a series on data provenance.

This report dissected a single company because it is the best documented. The method behind it is general: understanding a number means knowing its sample and knowing what it was calibrated against.

Data 02 takes on the opposite side: Who actually measures search engines and AI responses when both are giving off fewer and fewer clicks? And what happens to a panel when less and less traffic runs through a browser at all?

Related: Build it yourself on the question of what else you're buying with third-party components.