Enrich HubSpot contacts with LinkedIn: job title, company and job changes, every week
The job titles and companies in a CRM are right the day they're typed. Below, a Python script of 215 lines that keeps them as LinkedIn shows them. Run it every day: it reads the LinkedIn URL of each HubSpot contact, looks the profile up through Datacircle's API once a week, and writes the job title, company, city, state and country back into HubSpot. When the company in HubSpot is no longer among the person's current jobs, it also writes the date and the old company into two properties of its own, so a filter lists everyone who changed jobs.
Each check costs $1.25 per 1,000 profiles found, through Up2Data at its own price, and nothing for a profile that can't be reached: 1,000 contacts checked every week is $1.25 a week. The $5 you get at signup pays for the first 4,000 checks. HubSpot's free CRM is enough.
What you need
A Datacircle API key: sign up with your work email and it's on your dashboard, with the $5 credit. Python 3 and requests (pip install requests). Contacts whose LinkedIn URL property is filled: it's one of HubSpot's default contact properties, hs_linkedin_url in the API, and any way of writing the URL works, linkedin.com/in/someone included.
And a token for HubSpot's API, from a private app, which takes a super admin of the account, about two minutes:
- In HubSpot, Development, then Legacy apps in the left menu (HubSpot calls private apps legacy apps now, and still supports them).
- Create legacy app, then Private, and give it a name.
- On the Scopes tab, add four:
crm.objects.contacts.read,crm.objects.contacts.write,crm.schemas.contacts.readandcrm.schemas.contacts.write. The last two let the script add its three properties on its first run. - Create app, then on its Auth tab, Show token and copy it.
What it writes into HubSpot
Up2Data takes the profile's URL in a POST and answers the profile as it is on LinkedIn today: current_company is the person's main job, positions every job they list, with ended_at null for the ones they still hold, and location where they are. Each value goes into the HubSpot property that already holds it:
| HubSpot property | Internal name | Written from the profile |
|---|---|---|
| Job title | jobtitle | the title of the current job: the one at the company HubSpot already has, if they still work there, else their main job |
| Company name | company | that job's company, as LinkedIn names it |
| City, State/Region, Country/Region | city, state, country | location.city, region and country; a value LinkedIn doesn't give leaves yours as it is |
| LinkedIn checked on (new) | linkedin_checked_on | the day of the check, found or not: the next check is 7 days later |
| LinkedIn job change on (new) | linkedin_job_change_on | the day a job change was found |
| LinkedIn previous company (new) | linkedin_previous_company | the company HubSpot had before that change |
The first run adds the three new properties to your contacts, in the Contact information group. HubSpot keeps each property's history, so a title or company the script replaces is still on the contact's record. Every field of Up2Data's answer: LinkedIn Profile API in the docs.
What counts as a job change
The script compares the company HubSpot holds with every job the person holds today. Names are compared without their case, punctuation and legal endings, so "Acme Analytics, Inc." in HubSpot and "Acme Analytics" on LinkedIn are the same company. Many people hold several jobs at once: a side job you know them by, still held, isn't a change.
| Found | Written |
|---|---|
| a new company: HubSpot's company is no longer among their current jobs | the date, the old company, and the new company and title |
| no current job at all | the date and the old company; company and job title emptied |
| a job after none, on a contact checked before | the date, and the new company and title |
| a new title at the same company | the title only: a promotion isn't flagged, and the old title stays in HubSpot's history |
On a contact's first check, a company in HubSpot that the person no longer works at is a job change too: the first run over an old CRM finds everyone who has left the company HubSpot shows. Then, in HubSpot, filter your contacts on LinkedIn job change on in the last 7 days and save the view.
The whole file
Save this as enrich-hubspot-contacts-linkedin.py, or download the file, and run HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python enrich-hubspot-contacts-linkedin.py. Each run finds the contacts with a LinkedIn URL that were never checked, then those last checked 7 days ago or more, the oldest first, with HubSpot's search (pages of 200, at most 5 searches a second, as HubSpot asks). It checks them until Up2Data's daily limit and writes them back 100 at a time. A stop keeps every check made before it.
"""Keep your HubSpot contacts' job title, company and location as their LinkedIn profiles show them today, and flag who changed jobs.
https://datacircle.dev/blog/enrich-hubspot-contacts-linkedin
HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python enrich-hubspot-contacts-linkedin.py
HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python enrich-hubspot-contacts-linkedin.py --harvestapi (past Up2Data's daily limit, HarvestAPI)
Run it every day: it checks each contact with a LinkedIn URL once every 7 days, never-checked first, until Up2Data's daily limit (800 a day
per account), through Datacircle at $1.25 per 1,000 profiles found. HUBSPOT_TOKEN is a HubSpot private app's access token with the scopes
crm.objects.contacts.read, crm.objects.contacts.write, crm.schemas.contacts.read and crm.schemas.contacts.write.
"""
import datetime
import os
import re
import sys
import time
import requests
API = "https://api.datacircle.dev"
KEY = os.environ["DATACIRCLE_API_KEY"]
HUBSPOT = "https://api.hubapi.com"
HUBSPOT_HEADERS = {"Authorization": f"Bearer {os.environ['HUBSPOT_TOKEN']}"}
VERSION = "2026-09" # HubSpot's API version
EVERY_DAYS = 7 # how often each contact is checked
OVERWRITE = ["jobtitle", "city", "state", "country"] # written from LinkedIn; remove one to keep yours (the company is always written)
MAX_PER_RUN = 10000 # HubSpot's search gives at most 10,000 results
NEW_PROPERTIES = [ # added to your contacts on the first run
{"name": "linkedin_checked_on", "label": "LinkedIn checked on", "type": "date", "fieldType": "date"},
{"name": "linkedin_job_change_on", "label": "LinkedIn job change on", "type": "date", "fieldType": "date"},
{"name": "linkedin_previous_company", "label": "LinkedIn previous company", "type": "string", "fieldType": "text"},
]
READ = ["hs_linkedin_url", "jobtitle", "company", "linkedin_checked_on"]
SUFFIXES = {"inc", "llc", "ltd", "limited", "corp", "corporation", "co", "company", "gmbh", "plc", "sa", "ag", "bv", "lp", "llp"}
TODAY = datetime.datetime.now(datetime.timezone.utc).date() # Up2Data's daily limit resets at 00:00 UTC
def hubspot(method, path, body=None):
"""One HubSpot call. Its rate limits (429) and its own failures (5xx) are tried again after 10 s, then 30 s."""
for wait in (0, 10, 30):
time.sleep(wait)
answer = requests.request(method, f"{HUBSPOT}{path}", headers=HUBSPOT_HEADERS, json=body, timeout=60)
if answer.status_code != 429 and answer.status_code < 500:
break
try:
result = answer.json()
except ValueError:
result = {}
if answer.status_code in (401, 403):
sys.exit(f"HubSpot {answer.status_code}: {result.get('message')} Check HUBSPOT_TOKEN and its scopes (top of this file).")
return answer.status_code, result
def add_properties():
for prop in NEW_PROPERTIES:
status, _ = hubspot("GET", f"/crm/properties/{VERSION}/contacts/{prop['name']}")
if status == 404:
status, body = hubspot("POST", f"/crm/properties/{VERSION}/contacts", {**prop, "groupName": "contactinformation"})
if status not in (200, 201):
sys.exit(f"HubSpot {status} adding the property {prop['name']}: {body.get('message')}")
def due_contacts():
"""The contacts with a LinkedIn URL never checked, then those last checked EVERY_DAYS ago or more, the oldest first."""
cutoff = datetime.datetime.combine(TODAY - datetime.timedelta(days=EVERY_DAYS - 1), datetime.time(), datetime.timezone.utc)
has_url = {"propertyName": "hs_linkedin_url", "operator": "HAS_PROPERTY"}
searches = [[has_url, {"propertyName": "linkedin_checked_on", "operator": "NOT_HAS_PROPERTY"}],
[has_url, {"propertyName": "linkedin_checked_on", "operator": "LT", "value": str(int(cutoff.timestamp() * 1000))}]]
contacts = []
for filters in searches:
after = None
while len(contacts) < MAX_PER_RUN:
body = {"filterGroups": [{"filters": filters}], "properties": READ, "limit": 200,
"sorts": [{"propertyName": "linkedin_checked_on", "direction": "ASCENDING"}]}
if after:
body["after"] = after
status, page = hubspot("POST", f"/crm/objects/{VERSION}/contacts/search", body)
if status != 200:
sys.exit(f"HubSpot {status} searching contacts: {page.get('message')}")
contacts += page["results"]
after = page.get("paging", {}).get("next", {}).get("after")
if not after:
break
time.sleep(0.25) # HubSpot's search takes 5 requests a second
return contacts[:MAX_PER_RUN]
def text(value):
return (value or "").strip() # HarvestAPI's names and titles can end with a space
def place(city, state, country):
return {"city": text(city), "state": text(state), "country": text(country)}
def call(provider, url):
"""One lookup. A provider failing (502, 503, 504) or Up2Data's own rate limit isn't charged: wait, then try twice more."""
for wait in (0, 5, 15):
time.sleep(wait)
headers = {"Authorization": f"Token {KEY}", "X-Data-Provider": provider}
try:
if provider == "up2data":
answer = requests.post(f"{API}/v1/profiles/enrich", headers=headers, json={"url": url}, timeout=90)
else:
answer = requests.get(f"{API}/linkedin/profile", headers=headers, params={"url": url}, timeout=90)
except requests.RequestException as error:
return 0, {"error": str(error)} # no answer: the next run tries again
try:
body = answer.json()
except ValueError:
body = {} # a gateway's error page: the status says enough
rate_limited = answer.status_code == 429 and isinstance(body.get("error"), dict) # Datacircle's daily limit is a string
if not rate_limited and answer.status_code not in (502, 503, 504):
break
return answer.status_code, body
def check(url, provider):
"""One profile: (current jobs [(company, title)], location {city, state, country}, cost), "limit", "no profile", or None to try next run."""
status, body = call(provider, url)
if status == 401:
sys.exit("401: the Datacircle API key is wrong. It's on your dashboard.")
if status == 402:
sys.exit("402: your balance can't cover the call. Add funds on your dashboard, then run this again.")
if status == 429 and isinstance(body.get("error"), str):
return "limit"
if status == 400 or (provider == "up2data" and status == 422):
return "no profile" # not a profile URL, or a private or deleted profile: free
if provider == "up2data" and status == 200:
profile, cost = body["data"], body["datacircle_meta"]["cost_usd"]
main, where = profile.get("current_company") or {}, profile.get("location") or {}
jobs = [(text(main.get("name")), text(main.get("title")))] if main else []
jobs += [(text(job.get("company")), text(job.get("title"))) for job in profile.get("positions") or [] if not job.get("ended_at")]
return jobs, place(where.get("city"), where.get("region"), where.get("country")), cost
if provider == "harvestapi" and status == 200:
person, cost = body.get("element"), body["datacircle_meta"]["cost_usd"] # null when HarvestAPI can't find the profile, still billed
if not person:
return [], None, cost
where = (person.get("location") or {}).get("parsed") or {}
held = [job for job in person.get("experience") or [] if (job.get("endDate") or {}).get("text") == "Present"]
jobs = [(text(job.get("companyName")), text(job.get("position"))) for job in (person.get("currentPosition") or []) + held]
return jobs, place(where.get("city"), where.get("state"), where.get("country")), cost
reason = body.get("error", {}).get("message") if isinstance(body.get("error"), dict) else body.get("error", "")
print(f"{url}: {status or 'no answer'} {reason}, not charged; the next run tries again", file=sys.stderr)
return None
def plain(company):
"""A company's name to compare: "Acme, Inc." and "ACME" are the same company."""
words = re.sub(r"[^a-z0-9]+", " ", company.lower().replace("&", " and ")).split()
while words and words[-1] in SUFFIXES:
words.pop()
return " ".join(words)
def changes(contact, jobs, location):
"""The properties to write: LinkedIn's company, title and location, and the job change when the company in HubSpot is gone."""
old = contact["properties"]
company = (old.get("company") or "").strip()
props = {"linkedin_checked_on": TODAY.isoformat()}
if location is None: # no profile found: only the date of the check
return props, False
kept = next((job for job in jobs if company and plain(job[0]) == plain(company)), None) # the job HubSpot already knows, if it's still held
props["company"], props["jobtitle"] = kept or (jobs[0] if jobs else ("", ""))
props.update({name: value for name, value in location.items() if value}) # a place LinkedIn doesn't give leaves yours
for name in {"jobtitle", "city", "state", "country"} - set(OVERWRITE):
props.pop(name, None)
moved = bool(company and not kept) or bool(old.get("linkedin_checked_on") and not company and jobs) # a new company, gone, or a job after none
if moved:
props["linkedin_job_change_on"], props["linkedin_previous_company"] = TODAY.isoformat(), company
return props, moved
def save(updates):
for start in range(0, len(updates), 100): # HubSpot updates 100 contacts a call
status, body = hubspot("POST", f"/crm/objects/{VERSION}/contacts/batch/update", {"inputs": updates[start:start + 100]})
if status == 207:
for error in body.get("errors", []):
print(f"HubSpot: {error.get('message')}", file=sys.stderr)
elif status != 200:
print(f"HubSpot {status}: {body.get('message')}; these contacts are checked again next run", file=sys.stderr)
updates.clear()
def main(harvestapi):
add_properties()
contacts = due_contacts()
provider, found, checked, moved, spent, pending = "up2data", {}, 0, 0, 0.0, []
try:
for contact in contacts:
url = contact["properties"]["hs_linkedin_url"].strip()
if url not in found: # the same profile on two contacts is looked up once
result = check(url, provider)
if result == "limit" and harvestapi:
provider = "harvestapi"
result = check(url, provider)
if result == "limit":
print("Up2Data's daily limit: the next run, after 00:00 UTC, goes on (or add --harvestapi to go on now through HarvestAPI).")
break
if result is None:
continue
found[url] = ([], None, 0) if result == "no profile" else result
spent += found[url][2]
jobs, location, _ = found[url]
props, change = changes(contact, jobs, location)
pending.append({"id": contact["id"], "properties": props})
checked += 1
moved += change
if len(pending) == 100:
save(pending)
finally:
save(pending) # a stop keeps every check made
print(f"{checked} of {len(contacts)} due contacts checked, {moved} job changes (LinkedIn job change on: {TODAY}), ${spent:.5f} spent")
if __name__ == "__main__":
main("--harvestapi" in sys.argv[1:])What each answer does to the run:
| Answer | Written as | Cost |
|---|---|---|
| 200 | job title, company and location; a job change when there is one | $0.00125 |
| 422: private or deleted; 400: not a profile URL (a company page, say) | only LinkedIn checked on, so it's tried again in 7 days | free |
| 429 from Up2Data's own rate limit; 502, 503, 504 | tried again after 5 s, then 15 s; still failing, printed and left for the next run | free |
| 429: Datacircle's daily Up2Data limit | the run stops, and the next day's run goes on; with --harvestapi, the contact and the rest due go to HarvestAPI | free |
| 401: wrong key; 402: balance too low | the run stops; every check made before it is written to HubSpot | free |
| HubSpot 401 or 403: wrong token or a missing scope | the run stops before any check, with HubSpot's message | free |
| HubSpot 429 or 5xx | tried again after 10 s, then 30 s | free |
| HubSpot 207: a contact deleted while the script ran | HubSpot's message printed; the other contacts are written | as checked |
The same LinkedIn URL on two contacts is looked up once. One run takes at most 10,000 contacts, the most HubSpot's search gives for one query; the next run takes the rest.
Run it every day
Once a day is enough, at any hour: Up2Data's daily limit starts again at 00:00 UTC, and a contact HubSpot has just updated can take a few moments to reach its search. On a Mac or Linux, crontab -e and one line, here at 7:00 every morning:
0 7 * * * cd /path/to/folder && HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python3 enrich-hubspot-contacts-linkedin.py >> run.log 2>&1A contact whose LinkedIn URL you fill in today is checked by the next run, before the others. To check every 30 days instead of every 7, change EVERY_DAYS at the top of the file. To keep your own job titles or locations, remove them from OVERWRITE, next to it; the company is always written, since it's what the next check compares. The script makes a handful of HubSpot calls a run: a search per 200 contacts and a save per 100, far inside a private app's limits (100 calls every 10 seconds and 250,000 a day on Free and Starter).
What it costs
Every check is a live call, at the provider's price per profile found, taken from your balance with nothing added: $0.00125 through Up2Data, $0.0037 through HarvestAPI, and $0.0023 for a HarvestAPI lookup that finds nothing. HubSpot's API costs nothing more. Up2Data takes $1 a day per account (800 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC; HarvestAPI has no daily limit. So, every profile found:
| Contacts with a LinkedIn URL | Each checked | Through | Cost |
|---|---|---|---|
| 1,000 | every 7 days | Up2Data | $1.25 a week |
| 5,600 | every 7 days (800 a day) | Up2Data | $7.00 a week |
| 10,000 | about every 13 days (800 a day) | Up2Data | $7.00 a week |
| 10,000 | every 7 days, with --harvestapi | 800 through Up2Data, 9,200 through HarvestAPI | $35.04 a week |
Past 5,600 contacts, the script without the flag never checks more than Up2Data's 800 a day: each contact is simply checked less often, the oldest check first. With --harvestapi, every contact due is checked that day, those past the limit through HarvestAPI. Live. Each request goes to the provider and gets the profile as it is today.
HubSpot's own data enrichment
HubSpot enriches contacts and companies itself, from its knowledge base, Get started with data enrichment (updated July 10, 2026):
| HubSpot data enrichment | This script, through Datacircle | |
|---|---|---|
| Who gets it | a paid hub, Starter or above (Smart CRM and Revenue Hub: Professional or above) | any HubSpot account, the free CRM included |
| A contact is enriched from | first name, last name and work email | the contact's LinkedIn URL, its profile looked up live |
| A value you already have | kept, unless you choose to overwrite it | replaced by LinkedIn's (the old one stays in the property's history) |
| How often | "monthly when new data becomes available", once you turn it on | every 7 days, or what you set |
| Job changes | not mentioned | the date and the old company, in two properties |
| Price | comes with the subscription; uses no HubSpot Credits | $1.25 per 1,000 profiles found through Up2Data, $3.70 through HarvestAPI |
What HubSpot's enrichment does that the script doesn't:
- Enrich companies, from their domain, with industry and company size.
- Enrich a contact without a LinkedIn URL, from their name and work email.
- Run inside HubSpot, with no code, on new records as they come in.
The script gives you each contact as their LinkedIn profile shows them today, and who moved; the rest is HubSpot's.
How we tested it
We ran the file above as it is, its two hosts aside, on Python 3.9.6 and 3.12.14, against two stand-ins: one of HubSpot's API, built from its reference for version 2026-09 (the search with its filters, its sort, its pages of 200, its 10,000 results and 5 searches a second, saves of 100, a 207 for a contact deleted meanwhile, and a contact's update reaching the search 2 seconds late); one of ours, answering as its reference does, with people we made up and their jobs changed between runs. 19 contacts: a company spelled with "Inc." or "LLC", a side job, a URL with no https://, the same URL on two contacts, a private profile, a company page, a contact with no URL, one checked 3 days before, one deleted while the script ran.
- Week 1, with a daily limit of 8 found profiles: the first run checked 9 contacts and stopped at the limit; a second run that day checked none; the next day's checked 7 more, one after Up2Data's own rate limit and one after two 503s, and left one that answered 502 three times for the next run. Three job changes: two moves and a departure (company and title emptied). None for the "Inc." and "LLC" spellings or the side job; the shared URL was looked up once; HubSpot's 207 for the deleted contact was printed, the rest written.
- Week 2, with a limit of 3 and
--harvestapi: all 17 contacts due checked,$0.04395 spent, exactly 3 × $0.00125 + 9 × $0.0037 + 3 × $0.0023 (the company page, a 400, is free). Three job changes: a move, a side job left (the company is now the main job), a new job after none. A promotion wrote the new title, the old one in the history. HubSpot answered the first save with a 429; it went through 10 seconds later. - Week 3: the balance ran out after 4 paid checks; the run stopped on the 402 with those checks written to HubSpot.
- 10,250 contacts: the first run took 10,000 (50 searches, never past HubSpot's 5 a second, and 100 saves), the next run the other 250.
Both Pythons left the same contacts in the stand-in after each week. The first run on an account added the three properties; a token without the schema scopes stopped on HubSpot's 403 before anything was written. The same file, unchanged, against api.hubapi.com with a made-up token: HubSpot's 401, on paths that answer 404 with a made-up API version in them; against api.datacircle.dev with a wrong key: 401: the Datacircle API key is wrong, nothing written. We didn't run it on a real HubSpot account, having none to use: the test proves the script's logic and its calls, against the answers both APIs document.
The same weekly check for a CSV instead of a CRM: track job changes of your contacts. A whole CSV once: get LinkedIn profile data with Python; in a Google Sheet: LinkedIn profiles in Google Sheets. Both providers side by side: LinkedIn Profile API.
Questions
How do I enrich HubSpot contacts with LinkedIn data?
Read each contact's LinkedIn URL from HubSpot's API, look the profile up, and write the values back. With Datacircle: POST {"url": "<the contact's LinkedIn URL>"} to https://api.datacircle.dev/v1/profiles/enrich with your key and X-Data-Provider: up2data, then send data.current_company.title, data.current_company.name and data.location to HubSpot's batch update as jobtitle, company, city, state and country. The script on this page does it every day, each contact once a week, and flags who changed jobs.
Does HubSpot have a LinkedIn URL property?
Yes. LinkedIn URL is one of HubSpot's default contact properties, "the URL of the contact's LinkedIn page"; its internal name, the one the API uses, is hs_linkedin_url.
How much does it cost to keep 1,000 HubSpot contacts up to date every week?
$1.25 a week through Up2Data, for the profiles it finds; a private or deleted profile is free. Up2Data takes at most 800 a day per account, so up to 5,600 contacts can each be checked every week, for $7.00 a week. The $5 signup credit pays for the first 4,000 checks.
Does it work on HubSpot's free CRM?
Yes: HubSpot's contacts search and batch update APIs are on every tier, Free included, and a private app on Free or Starter gets 100 calls every 10 seconds and 250,000 a day. Creating the private app takes a super admin of the HubSpot account. HubSpot's own data enrichment needs a paid hub, Starter or above.
Is each request live, or cached?
Live. Each request goes to the provider and gets the profile as it is today.
Is there a daily limit?
Up2Data takes $1 a day per account (800 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC; HarvestAPI has no daily limit.
Sign up at datacircle.dev with your work email: a $5 credit, that's 4,000 LinkedIn profiles at $1.25 per 1,000. Free: 10M+ U.S. B2B leads, as a flat file. Download it at datacircle.dev.
Sign up