Artificially generatedThe Super Data Collectors
How Similarweb records the browsing behavior of the entire world.
An Israeli startup provides visitor numbers for one hundred million websites without operating a counter on a single one of them. This is not magic and not a scandal, but a very cleverly built machine. This report takes it apart: which four sources feed into it, how three million browsers become a statement about one hundred million websites, where the panel comes from, who built the company, and how accurate the result is in the end.
How does someone know how many visitors a third-party website has?
Every second pitch deck contains a Similarweb number. Competitive analyses, due diligence, media planning, hedge fund models: somewhere in there sits a bar chart with monthly visits to a website to which the author has no access whatsoever. And it works surprisingly well.
The question behind it is a nice technical question, and it has a clean answer. The number is not a measurement, it is an extrapolation. Just like an election forecast: a sample is observed, corrected, calibrated against known truths, and then stretched to the total population.
Anyone who understands an election forecast understands Similarweb. It's the same four work steps, the same sources of error, the same kind of caution when reading the result. The only difference is the sample: instead of a thousand people called, it's millions of browsers.
This report goes through the machine from the beginning. First the four inputs, as the company itself describes them. Then the calculation method. Then the origin of the panel, which is its own and quite interesting story. After that the company, the numbers, and the accuracy.
First the scale: Similarweb is not a special case in this industry, but the best-documented example. What is written here applies in principle to every provider of market data from behavioral measurement.
What flows in, and what each source is good for
The middle column of the following table is not a summary, but the translation of the wording from the manufacturer's page Our Data. The right column explains the role in the calculation process. It is important that the four sources do not do the same thing: one delivers truth, one delivers behavior, one delivers structure, one fills gaps.
Order as on the manufacturer's page. How strongly each source contributes to a single number is not published.
| Source | Own description | Role in calculation process |
|---|---|---|
| Website and app operators | Directly measured data from their own analytics software (e.g. Google Analytics) from millions of websites and apps. | The calibration standard. Operators voluntarily provide their real numbers in exchange for comparative values. Only here does Similarweb know the truth, and only against this can the model be adjusted. |
| Contributors Network | Anonymous traffic data collected from Similarweb products installed on millions of devices worldwide. | The panel. Browser extensions and apps that record which addresses are accessed. The actual raw material: only here does behavior emerge across third-party websites. |
| Public data | Publicly available data (e.g. Wikipedia, census data), algorithmically captured and indexed. | The map. Crawlers provide structure, categories, search result pages, app store positions, and population numbers. Says little about visitors, but a lot about what to compare them against. |
| Partnerships | Pre-analyzed data from global partners such as DSPs, internet providers, measurement service providers, and credit agencies. | The gap filler. Purchased movement data from sources that the own panel does not reach: other age groups, other countries, other devices. Greatly expanded after 2018. |
How three million browsers become one hundred million websites
This is the core of the whole thing, and it is surprisingly comprehensible. Four steps, each with its own error, which the next step partially captures again.
The second peculiarity follows from step 02. A panel consisting of browser extensions sees the world through a desktop browser. Everything that bypasses it, it sees poorly: apps, embedded views in social networks, and increasingly answers provided by an AI system without anyone accessing a website. That's why the company has been buying partner data and entire companies for years. The panel is not the product, it is the calibration basis that must be constantly updated.
Millions of people contribute measurements without that being their goal
Now for the most difficult ingredient. A panel of this size doesn't arise from people voluntarily signing up for market research. It arises from a product being useful and recording on the side.
This isn't a Similarweb invention, it's the basic model of behavioral measurement since TV ratings have existed. What's new is the scale: instead of a few thousand households, it's millions of devices, and the compensation isn't cash but a free tool.
The origin of this panel has been documented over ten years by security researchers, universities, and an ad blocker manufacturer. The following timeline is the sober summary of these works.
Only incidents with published technical analysis. The reach figures cannot be added, they partially overlap.
| Date | Finding | Reach | Who |
|---|---|---|---|
| 03/2016 | 42 Chrome extensions contain the library upalytics and send every visited address, every search query, and even addresses from internal corporate networks to nine domains, all registered through an anonymization service. Among them similarsites.com. The same library is embedded in the Similarweb extension itself. | 8.0M | Michael Weissbacher, Northeastern University |
| 01/2017 | Similarweb acquires the design extension Stylish, with which users change the appearance of third-party websites. Starting with the same version, the Chrome version begins recording. | ~2.0M | Acquisition, public |
| 07/2018 | Stylish sends every complete address along with a permanent identifier to api.userstyles.org, double base64-encoded. Complete address means in practice also: login tokens, password reset links, search results. Google and Mozilla remove the extension from their offerings within two days, a revised version returns in August. | ~2.0M | Robert Heaton |
| 07/2018 | Nine tools from a newly founded Delaware company Big Star Labs send every visited page to their own servers. AdGuard determines the data format resembles that of Stylish, but explicitly calls a connection unconfirmed. | 11.0M | AdGuard |
| 12/2025 | The Similarweb extension reads conversations from ChatGPT, Claude, Gemini, and Perplexity, via a dynamically loaded configuration file with its own evaluation logic for each provider. The capability was added with an update in May 2025. The researcher calls the pattern Prompt Poaching. On January 1, 2026, Similarweb adds an explicit consent dialog for this. | 1.0M | John Tuckner, Secure Annex |
| 02/2026 | Stylish transmits addresses and AI conversations through five-layer obfuscation: URL encoding, double base64, column swapping, AES-256-CBC with key embedded in source code, finally base64 again. | ~2.0M | James Arnott |
| 05/2026 | Systematic scan of 287 extensions with a combined 37.4M users, roughly one percent of all Chrome users worldwide. Similarweb appears as the most frequent recipient via multiple pathways. The report also attributes Big Star Labs (3.7M users) to the company. | 37.4M | Q Continuum |
From 2018 onwards, the situation changes. Google and Mozilla tighten their rules, the European General Data Protection Regulation takes effect, and Similarweb audibly restructures: more purchased partner data, less dependence on the Store. In the 2021 agreement with security company Check Point, it even explicitly states that Stylish data is not part of the exchange. In its securities filings, the company still lists dependence on the Contributory Network as its own risk factor: platform rules change frequently and are enforced inconsistently.
How many extensions does Similarweb operate today?
The obvious answer is two: the traffic extension and the sales extension, both bearing the name. Instead, on August 7, 2026, we extracted the publisher IDs from the Chrome Web Store, meaning we compared accounts, not product names. This adds two more entries.
Stylish still appears under the publisher account similarweb-ltd and with two million users is the largest item. Similar Sites runs under similargroup, the old company name. Neither carries the Similarweb name in the Store listing.
User numbers as shown by the Chrome Web Store, retrieved on August 7, 2026. The Store rounds to even increments.
| Extension | Publisher in Store | Bears the name? | Users |
|---|---|---|---|
| Stylish: Custom themes for any website | similarweb-ltd | no | 2,000,000 |
| Similarweb: Website Traffic, AI Traffic & SEO Checker | Similarweb | yes | 1,000,000 |
| Similar Sites: Discover Related Websites | similargroup | no | 300,000 |
| Similarweb Sales Extension | Similarweb | yes | 20,000 |
| Chrome total | 3,320,000 |
The answer is in the terms, and it is surprisingly open
For the question of what data actually flows, you don't need a security researcher. You just have to read the right file, and that's the small difficulty: Similarweb's privacy policy explicitly does not apply to the browser extension. It says so in its own sentence.
The declaration that applies to the extension is on a different domain: similarsites.com. There it is stated very clearly who is responsible and what is collected. This is not a leak and not a discovery, this is the text that is agreed to upon installation.
The last line in the table is the most economically important: the result may be passed on to business customers. This is exactly the product that Similarweb sells.
Verbatim from the privacy policy at similarsites.com, accessed on August 7, 2026.
| Category | Own wording |
|---|---|
| Responsible | "Similarweb Ltd., incorporated under the laws of the State of Israel, is the controller" |
| Addresses | "URLs accessed or visited, pages on which advertising was seen or clicked, advertising URL" |
| History | "Clickstream data, search history and information about interactions" |
| Search | "Data from search results pages (search term, order or rank of results)" |
| AI inputs | "Prompts, queries, content and other inputs that you enter into or submit to certain artificial intelligence tools" |
| Sharing | "We may 'sell' or 'share' these categories to business customers so that our business customers can better understand consumer behavior" |
From jewelry store to stock exchange: Or Offer
Or Offer, born in 1983, grows up in Tzur Hadassah near Jerusalem, as a computer kid with parents who design jewelry and struggle with the uncertainty of this profession. After school, intelligence unit 8200 wants him as a programmer. He declines and goes to Oketz, the canine unit, among other things to a post in the Gaza Strip. About this time he later says it taught him to stand up to experienced people as a young person.
Afterwards he studies business administration part-time at IDC Herzliya and builds two jewelry stores, in Zichron Ya'akov and Netanya, plus an online shop. The chain is called Maya Offer, after his mother. His parents push him towards high-tech.
The trigger is a name: a supplier mentions the designer David Yurman, who combines stones with silver. Offer wants to find comparable products online and realizes that there is no tool for this. Similarweb is created in 2007 as a browser extension that suggests similar websites to a website. Hence the name, and hence also the panel: data collection is originally the byproduct of a search function.
He is 23 and works for over a year without salary. In 2007 an important collaborator leaves the project because he doesn't believe in the product. At 24 Offer thinks about quitting and stays: "I had nothing else to do, so I continued."
In 2009 Yossi Vardi gives him one million dollars, for about five percent. The company lives on this for five years, with seven or eight people. In 2011 comes the pivot that decides everything: away from the consumer product, towards an analysis platform for companies and investors. The consumer product doesn't disappear in the process, it changes roles. From now on it is no longer the product, it is the measuring point.
Offer was sole founder and considers this an advantage: "When there are multiple founders, each needs a role to feel important. Precisely because I was alone, that saved me a lot of politics."
In 2014 Naspers leads a round of 18 million dollars. In May 2021 the company goes public on the New York Stock Exchange, valued at 1.6 billion dollars, and raises around 180 million dollars. Headquarters is still Givatayim near Tel Aviv, plus New York and London.
The growth rates of over 40 percent occurred in the years around the IPO. Since 2023 growth has been stable in the low teens.
| Year | Revenue (million $) | Growth |
|---|---|---|
| 2019 | 70.6 | – |
| 2020 | 93.5 | +32.4 % |
| 2021 | 137.7 | +47.3 % |
| 2022 | 193.2 | +40.4 % |
| 2023 | 218.0 | +12.8 % |
| 2024 | 249.9 | +14.6 % |
| 2025 | 282.6 | +13.1 % |
Breadth was acquired
The company's own panel measures web traffic in the browser. Everything beyond that, app data, advertising data, search engine rankings, came through acquisitions. Eight documented purchases in ten years, purchase prices almost never published. The list reads like a map of the blind spots that a browser panel has.
| Year | Company | Which gap was closed with it |
|---|---|---|
| 2014 | TapDog | Early-stage startup, shares and cash |
| 2015 | Swayy | Content discovery |
| 2015 | Quettra | Mobile analysis, i.e. the first blind spot of the browser panel |
| 2017 | Stylish | Around two million browsers at once, i.e. pure panel size |
| 2021 | Embee Mobile | Mobile panel data |
| 2022 | Rank Ranger | Search engine rankings and interfaces |
| 2024 | Admetricks | Advertising data |
| 2024 | 42matters | App data |
| 2025 | The Search Monitor | Paid search, brand protection, affiliate program control |
The raw material becomes training material
On May 13, 2026, the board of directors opened the search for a new CEO. Or Offer is leaving by mid-2027, after almost exactly twenty years. "Similarweb was my life's work," he says, it is "the right moment." Board chairman Harel Beit-On calls the result "a unique data asset."
The market capitalization in May 2026 was around 270 million dollars, a good 56 percent below the beginning of the year and far below the 1.6 billion of the IPO. Revenue is growing, the stock price is not. That is the framework in which to read the next step.
In the first quarter of 2026, Similarweb signed a seven-figure contract for training data for a large language model, with an existing major customer. A second is announced. Together with other multi-year deals, the company cites 47 million dollars in total contract value.
This creates a remarkable loop. The panel records how people talk to AI systems. The analysis of this is sold to the manufacturers of such systems. And a third product sells brands the observation of how these systems recommend their products. The same raw material, three markets.
For the reader, the interesting thing about this is the direction: behavioral data is currently in the process of becoming a raw material for models rather than an analysis tool. Similarweb is an early, well-documented case of this.
The error is large, and it doesn't always point in the same direction
There are several independent comparisons against real Analytics access. They contradict each other in sign, and that is precisely the result. Anyone expecting a single error figure has misunderstood the matter: the error depends on how large the site is, what type of site it is, and how well Analytics itself measures on that site.
All against Google Analytics as reference. Different samples, different years, different signs.
| Study | Sample | Finding |
|---|---|---|
| Omniconvert, 2026 | 1,787 shops | Sessions around 94 percent too high. Error systematic, not random. Large sites significantly more accurate than small ones. |
| PLOS One, 2022 | 86 websites | Visits 19.4 percent too low, unique visitors 38.7 percent too low, bounce rate 25.2 percent too high. |
| SparkToro | Provider comparison | For sites with 5,000+ users most frequently within ±30 percent and thus better than all competitors. Below 5,000 users worse. |
Four sentences to get the picture right
First: nothing is being accused here. There is no case in which a legal violation has been proven to an authority. This report describes a mechanism, it does not accuse.
Second: Similarweb contradicts the 2016 presentation. The company has stated that these are "repeated foreign reports from 2016 that were already baseless at the time and were refuted many years ago." This sentence stands here as it was made.
Third: the company has visibly restructured. After 2018, away from dependence on the Store, toward purchased partner data, with explicit exclusion of Stylish data in the Check Point agreement and a consent dialog for AI content since January 2026.
Fourth: this is an industry model, not an isolated case. Behavioral panels from consumer products have existed since television ratings. The 2026 scan found 287 extensions, not four. Similarweb is the best-documented example here because it is publicly traded and therefore must be accountable.
And the point at which the whole thing dissolves: there is no method to estimate third-party traffic without observing someone somewhere. The alternative to a panel is not a better panel, but no number at all. Anyone who uses the number uses the method with it.
For your analyses
Use Similarweb for ratios and trends, not for absolute numbers in a calculation on which something depends. The documented error ranges from half to double and reverses sign depending on site type.
For small sites, be doubly cautious
Below about 5,000 monthly users it becomes unreliable, for a structural reason: there are too few calibration points and too few panel hits there.
For your own site
Your numbers are already in the panel, and your competition can read them. This is not an attack, but the normal case. Plan with it.
For your work devices
A browser extension sees everything that happens in the tab. On a computer with customer data, contracts, or AI conversations in the browser, every installed extension is a decision about confidentiality. Counting them once and checking the publisher for each one is a matter of ten minutes.
- Or Offer, extensive portrait interview (Calcalist/CTech). Childhood, Oketz, the jewelry chain, the David Yurman moment, the years without salary, Vardi's million, the pivot in 2011. The source for almost all quotes in section 07. calcalistech.com
- Or Offer in conversation with Forbes, February 2017. Entrepreneurship, crises, why a pivot is "like a divorce." forbes.com
- Or Offer on his departure, May 2026 (CTech). "Similarweb was my life's work," plus the current business figures and the stock price decline. calcalistech.com
- Original announcement on succession, May 13, 2026. Wording from Or Offer and board chairman Harel Beit-On. ir.similarweb.com
- Marta Sulkiewicz, VP Emerging Solutions, on the AI business (MediaPost, May 2026). On the training data contract and what brands are asking today: "Every day we are asked if we can see purchase completions in AI agents and bots." mediapost.com
- Michael Weissbacher, Northeastern University, on the 2016 investigation. The initial description of the panel from a security research perspective. mweissbacher.com
- John Tuckner, Secure Annex, on Prompt Poaching. "Prompt Poaching has arrived to siphon off your most sensitive conversations, and browser extensions are the way there." secureannex.com
- Quarterly reports with Or Offer in original audio. All transcripts and recordings of the conference calls. ir.similarweb.com
Everything this report is based on
Fourteen sources, each with what it contributes and the full address. They are intentionally set large and expanded: those who don't need them scroll past, those who want to verify shouldn't have to search.
- Similarweb: Our DataThe four data sources in the manufacturer's own words, plus the size specifications: over 100 million websites, over 4 million apps, 235 million product articles, 10 billion signals and 2 TB per day, 200 data scientists and 50 PhDs.similarweb.com/corp/ourdata/
- Similarweb: Privacy PolicyContains the crucial sentence that this policy explicitly does not apply to the app and browser extension. This begins the trail to the policy that actually applies.similarweb.com/corp/legal/privacy-policy/
- SimilarSites: Privacy PolicyThe policy that actually applies to the extension. Names Similarweb Ltd. as the controller and lists complete URLs, clickstream, search history and AI inputs, plus the sale to business customers. The basis for Fig. 04.similarsites.com/privacy-policy
- Similarweb: Form 20-F, Fiscal Year 2021 (SEC)The risk disclosure in which the company itself identifies the Contributory Network's dependence on extensions and apps in third-party stores as a business risk. The US Securities and Exchange Commission server rejects automated requests, but the document is normally accessible in the browser.sec.gov/Archives/edgar/data/1842731/000184273122000013/smwb-20211231.htm
- Similarweb Investor Relations: Quarterly ResultsRevenue 2025, customer count, customers with annual revenue of $100,000 or more, the first quarter of 2026, and the seven-figure contract for training data.ir.similarweb.com/financials/quarterly-results
- Michael Weissbacher, Northeastern University: These Chrome extensions spy on 8 million users (March 2016)The initial description. 42 extensions with the upalytics library, 8 million installations combined, nine recipient domains registered through an anonymization service. The starting point of the entire chain of evidence.mweissbacher.com/2016/03/31/these-chrome-extensions-spy-on-8-million-users/
- Ex-Ray: Detection of History-Leaking Browser Extensions (ACSAC 2017)The scientific continuation from Northeastern University and University College London. 10,691 extensions examined, 212 identified as history-leaking, solely from the pattern of network traffic.mweissbacher.com/publications/acsac_exray.pdf
- Robert Heaton: Stylish steals all your internet history (July 2, 2018)The technical analysis that led to removal from the stores. Complete addresses including persistent identifier to api.userstyles.org, double base64-wrapped, linkable to the account via the session cookie.robertheaton.com/2018/07/02/stylish-browser-extension-steals-your-internet-history/
- AdGuard: Big Star Labs spyware campaign affects over 11.000.000 people (July 2018)Nine tools from a Delaware company, recipient domains, privacy policies as image files. Contains the explicit caveat that the suspected connection to Similarweb is not confirmed. That's exactly why it's also listed here as unconfirmed.adguard.com/en/blog/big-star-labs-spyware.html
- John Tuckner, Secure Annex: Prompt poaching runs rampant in extensions (December 2025)The original analysis on intercepting AI conversations: dynamically loaded configuration file with custom evaluation logic per provider, for ChatGPT, Claude, Gemini and Perplexity. Coins the term Prompt Poaching.secureannex.com/blog/prompt-poaching/
- Q Continuum: Report on spying browser extensions (2026)The largest systematic scan to date: 287 extensions with 37.4 million users, 930 CPU-days of measurement time, methods and individual findings disclosed. Also attributes Big Star Labs to the company, which Similarweb has not confirmed.github.com/qcontinuum1/spying-extensions
- Consumer Rights Wiki: Entry SimilarWebThe continuous timeline of incidents through May 2026, including James Arnott's analysis of the five-level obfuscation and the classification of both extensions as confirmed exfiltrators.consumerrights.wiki/w/SimilarWeb
- Globes: Similarweb's controversial route to Wall Street (2021)The most comprehensive journalistic investigation. Contains the company's statement verbatim and the agreement with Check Point, in which Stylish data is explicitly excluded.en.globes.co.il/en/article-similarwebs-controversial-route-to-wall-street-1001376912
- Accuracy studies: Omniconvert (2026) and PLOS One (2022)The two comparisons against real analytics access from section 09. Omniconvert across 1,787 shops, PLOS One across 86 websites, with opposite signs. Both are linked because only the contradiction carries the statement.omniconvert.com/blog/we-analyzed-1787-ecommerce-websites-similarweb-google-analytics-thats-we-learned/ journals.plos.org/plosone/article?id=10.1371/journal.pone.0268212
McGrinsey's own work: The user numbers and publisher accounts of the four Chrome extensions in Fig. 03 were read directly from the store entries on August 7, 2026. The Chrome Web Store rounds user numbers to smooth levels, so the sum of 3,320,000 is correspondingly rough. The calculation method in section 03 is a reconstruction from the published methodology and the academic literature on panel extrapolation; the company does not publish it in detail. All translations are ours. Where sources contradict each other, the contradiction is stated in the text.
This report dissected a single company because it is the best documented. The method behind it is general: understanding a number means knowing its sample and knowing what it was calibrated against.
Data 02 takes on the opposite side: Who actually measures search engines and AI responses when both are giving off fewer and fewer clicks? And what happens to a panel when less and less traffic runs through a browser at all?
Related: Build it yourself on the question of what else you're buying with third-party components.


