Most B2B data comes through an API: you ask for one record, you pay for one record, you wait for one record. We do that too, live, with no markup. But we also hand out flat files, and our founder, Wayne, wanted the reason written down: "the benefit of flat files."
Here it is, measured on our free file on the day we wrote this.
What's in the file
Free: 10M+ U.S. B2B leads, as a flat file. Download it at datacircle.dev.
10M+ U.S. B2B: 9.5M people and 1.75M companies, two Parquet files in one 2.3 GB zip.
One table per file. The people file has 9,521,154 rows and 23 columns (name, headline, LinkedIn URL, city and state, current job title, start date and company, email, mobile phone, when it was updated). The companies file has 1,748,139 rows and 19 columns (name, industry, headcount range, headquarters, founded, website). Every column is typed, and a person points to their company by CURRENT_JOB_COMPANY_LINKEDIN_ID.
1. You pay once, then every question is free
Pulling 9.5M profiles through an API, one call each, at $1.25 per 1,000 would be $11,901, and you'd pay again the next time you change your mind about who you want. The file is one download. After that, a question costs nothing: not the first one, not the hundredth.
2. It answers in under a second
DuckDB 1.5.4 on an Apple M2 Max laptop, the two files on its disk:
import duckdb
duckdb.sql("create view people as select * from 'us_smb_mid_market_persons.parquet'")
duckdb.sql("create view companies as select * from 'us_smb_mid_market_companies.parquet'")
| The question | Time | The answer |
|---|---|---|
| People per state | 0.01 s | California 1,157,818, Texas 876,796, Florida 665,456, New York 635,110, Illinois 371,104 |
| How many have an email, a mobile phone, either | 0.06 s | 2,790,286, 4,016,020, 5,182,067 |
| Owners at Texas companies of 11 to 50 people, by industry (a join of the two files) | 0.13 s | Construction 887, no industry given 454, Retail 331, Real Estate 307, Individual and Family Services 305 |
The third one is a join across 9.5M people and 1.75M companies:
select c.LINKEDIN_INDUSTRY, count(*)
from people p join companies c on c.LINKEDIN_ID = p.CURRENT_JOB_COMPANY_LINKEDIN_ID
where c.EMPLOYEE_COUNT_RANGE = '11-50' and c.HQ_STATE_CODE = 'TX' and p.CURRENT_JOB_TITLE ilike '%owner%'
group by 1 order by 2 desc
No pagination, no rate limit, no API key. Through an API, the same answer is thousands of calls.
3. Parquet reads only what you ask for
Parquet stores a file column by column, compressed. The people file holds 5.25 GB of data in 1.71 GB. And a query reads only the columns it names: STATE_CODE is 6.6 MB of those 1.71 GB, so "people per state" reads 6.6 MB, not the whole file. That's why it takes 0.01 s. count(*) doesn't even read a column: the row count is written in the file's footer.
The heaviest columns are the long texts (ABOUT, 523 MB; PROFILE_PIC_URL, 324 MB). A query that doesn't name them never pays for them.
4. It's yours
- Nothing leaves your machine. Join your own CRM export with
join read_csv('my_accounts.csv')and nobody sees who your accounts are. - The same file gives the same answer. Every file comes with its
sha256, so you can check your download and know two people ran the same query on the same data. - You can compare months. We publish a new 50M+ file every month. When October's came out smaller, comparing the two files gave the answer:
September's had 50.5M people and 6.8M companies (11.9 GB); October's has 43.6M people and 6.7M companies (9.6 GB). Same columns, same filters: fewer people in this month's source pass them.
- Any tool reads it. DuckDB, pandas, Polars, Spark, ClickHouse:
import pandas as pd
people = pd.read_parquet("us_smb_mid_market_persons.parquet", columns=["FULL_NAME", "CURRENT_JOB_TITLE", "STATE_CODE"])
What a file can't do
A file is a snapshot. In this one, UPDATED_AT runs from August 2024 to September 2026. When you need someone as they are today, that's the API's job:
Live. Each request goes to the provider and gets the profile as it is today.
So the two work together: find who you want in the file, then pull the live profile of the ones you picked. That's the co-op:
Step 1: Query your favorite B2B data APIs through us. Same request, same price, no markup.
Step 2: You're DONE. Every morning, you get the flat file of your data plus everyone else's.
Every profile anyone pulls through us goes into a co-op file, a new one every morning, one flat Parquet table per provider, read with the same three lines of DuckDB:
Add $50 to your account: you get $50 of API PLUS the flat file. Or invite 3 people who sign up.
Get it
Sign up at datacircle.dev with your work email: a $5 credit, that's 4,000 LinkedIn profiles at $1.25 per 1,000.
Then download the 10M+ file from your dashboard (two minutes on our connection, see how R2 makes that free for us in "How R2's zero egress lets us give away a 9.6 GB file"), and every question after that is free. For questions an API can't ask at all, read "Why SQL is superior when building your list".