Datacircle

Enrich HubSpot contacts with LinkedIn: job title, company and job changes, every week

The job titles and companies in a CRM are right the day they're typed. Below, a Python script of 215 lines that keeps them as LinkedIn shows them. Run it every day: it reads the LinkedIn URL of each HubSpot contact, looks the profile up through Datacircle's API once a week, and writes the job title, company, city, state and country back into HubSpot. When the company in HubSpot is no longer among the person's current jobs, it also writes the date and the old company into two properties of its own, so a filter lists everyone who changed jobs.

Each check costs $1.25 per 1,000 profiles found, through Up2Data at its own price, and nothing for a profile that can't be reached: 1,000 contacts checked every week is $1.25 a week. The $5 you get at signup pays for the first 4,000 checks. HubSpot's free CRM is enough.

What you need

A Datacircle API key: sign up with your work email and it's on your dashboard, with the $5 credit. Python 3 and requests (pip install requests). Contacts whose LinkedIn URL property is filled: it's one of HubSpot's default contact properties, hs_linkedin_url in the API, and any way of writing the URL works, linkedin.com/in/someone included.

And a token for HubSpot's API, from a private app, which takes a super admin of the account, about two minutes:

  1. In HubSpot, Development, then Legacy apps in the left menu (HubSpot calls private apps legacy apps now, and still supports them).
  2. Create legacy app, then Private, and give it a name.
  3. On the Scopes tab, add four: crm.objects.contacts.read, crm.objects.contacts.write, crm.schemas.contacts.read and crm.schemas.contacts.write. The last two let the script add its three properties on its first run.
  4. Create app, then on its Auth tab, Show token and copy it.

What it writes into HubSpot

Up2Data takes the profile's URL in a POST and answers the profile as it is on LinkedIn today: current_company is the person's main job, positions every job they list, with ended_at null for the ones they still hold, and location where they are. Each value goes into the HubSpot property that already holds it:

The HubSpot contact properties the script writes, and where each value comes from
HubSpot propertyInternal nameWritten from the profile
Job titlejobtitlethe title of the current job: the one at the company HubSpot already has, if they still work there, else their main job
Company namecompanythat job's company, as LinkedIn names it
City, State/Region, Country/Regioncity, state, countrylocation.city, region and country; a value LinkedIn doesn't give leaves yours as it is
LinkedIn checked on (new)linkedin_checked_onthe day of the check, found or not: the next check is 7 days later
LinkedIn job change on (new)linkedin_job_change_onthe day a job change was found
LinkedIn previous company (new)linkedin_previous_companythe company HubSpot had before that change

The first run adds the three new properties to your contacts, in the Contact information group. HubSpot keeps each property's history, so a title or company the script replaces is still on the contact's record. Every field of Up2Data's answer: LinkedIn Profile API in the docs.

What counts as a job change

The script compares the company HubSpot holds with every job the person holds today. Names are compared without their case, punctuation and legal endings, so "Acme Analytics, Inc." in HubSpot and "Acme Analytics" on LinkedIn are the same company. Many people hold several jobs at once: a side job you know them by, still held, isn't a change.

What the script writes to LinkedIn job change on, and when
FoundWritten
a new company: HubSpot's company is no longer among their current jobsthe date, the old company, and the new company and title
no current job at allthe date and the old company; company and job title emptied
a job after none, on a contact checked beforethe date, and the new company and title
a new title at the same companythe title only: a promotion isn't flagged, and the old title stays in HubSpot's history

On a contact's first check, a company in HubSpot that the person no longer works at is a job change too: the first run over an old CRM finds everyone who has left the company HubSpot shows. Then, in HubSpot, filter your contacts on LinkedIn job change on in the last 7 days and save the view.

The whole file

Save this as enrich-hubspot-contacts-linkedin.py, or download the file, and run HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python enrich-hubspot-contacts-linkedin.py. Each run finds the contacts with a LinkedIn URL that were never checked, then those last checked 7 days ago or more, the oldest first, with HubSpot's search (pages of 200, at most 5 searches a second, as HubSpot asks). It checks them until Up2Data's daily limit and writes them back 100 at a time. A stop keeps every check made before it.

"""Keep your HubSpot contacts' job title, company and location as their LinkedIn profiles show them today, and flag who changed jobs.
https://datacircle.dev/blog/enrich-hubspot-contacts-linkedin

HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python enrich-hubspot-contacts-linkedin.py
HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python enrich-hubspot-contacts-linkedin.py --harvestapi   (past Up2Data's daily limit, HarvestAPI)
Run it every day: it checks each contact with a LinkedIn URL once every 7 days, never-checked first, until Up2Data's daily limit (800 a day
per account), through Datacircle at $1.25 per 1,000 profiles found. HUBSPOT_TOKEN is a HubSpot private app's access token with the scopes
crm.objects.contacts.read, crm.objects.contacts.write, crm.schemas.contacts.read and crm.schemas.contacts.write.
"""
import datetime
import os
import re
import sys
import time

import requests

API = "https://api.datacircle.dev"
KEY = os.environ["DATACIRCLE_API_KEY"]
HUBSPOT = "https://api.hubapi.com"
HUBSPOT_HEADERS = {"Authorization": f"Bearer {os.environ['HUBSPOT_TOKEN']}"}
VERSION = "2026-09"  # HubSpot's API version
EVERY_DAYS = 7  # how often each contact is checked
OVERWRITE = ["jobtitle", "city", "state", "country"]  # written from LinkedIn; remove one to keep yours (the company is always written)
MAX_PER_RUN = 10000  # HubSpot's search gives at most 10,000 results
NEW_PROPERTIES = [  # added to your contacts on the first run
    {"name": "linkedin_checked_on", "label": "LinkedIn checked on", "type": "date", "fieldType": "date"},
    {"name": "linkedin_job_change_on", "label": "LinkedIn job change on", "type": "date", "fieldType": "date"},
    {"name": "linkedin_previous_company", "label": "LinkedIn previous company", "type": "string", "fieldType": "text"},
]
READ = ["hs_linkedin_url", "jobtitle", "company", "linkedin_checked_on"]
SUFFIXES = {"inc", "llc", "ltd", "limited", "corp", "corporation", "co", "company", "gmbh", "plc", "sa", "ag", "bv", "lp", "llp"}
TODAY = datetime.datetime.now(datetime.timezone.utc).date()  # Up2Data's daily limit resets at 00:00 UTC


def hubspot(method, path, body=None):
    """One HubSpot call. Its rate limits (429) and its own failures (5xx) are tried again after 10 s, then 30 s."""
    for wait in (0, 10, 30):
        time.sleep(wait)
        answer = requests.request(method, f"{HUBSPOT}{path}", headers=HUBSPOT_HEADERS, json=body, timeout=60)
        if answer.status_code != 429 and answer.status_code < 500:
            break
    try:
        result = answer.json()
    except ValueError:
        result = {}
    if answer.status_code in (401, 403):
        sys.exit(f"HubSpot {answer.status_code}: {result.get('message')} Check HUBSPOT_TOKEN and its scopes (top of this file).")
    return answer.status_code, result


def add_properties():
    for prop in NEW_PROPERTIES:
        status, _ = hubspot("GET", f"/crm/properties/{VERSION}/contacts/{prop['name']}")
        if status == 404:
            status, body = hubspot("POST", f"/crm/properties/{VERSION}/contacts", {**prop, "groupName": "contactinformation"})
            if status not in (200, 201):
                sys.exit(f"HubSpot {status} adding the property {prop['name']}: {body.get('message')}")


def due_contacts():
    """The contacts with a LinkedIn URL never checked, then those last checked EVERY_DAYS ago or more, the oldest first."""
    cutoff = datetime.datetime.combine(TODAY - datetime.timedelta(days=EVERY_DAYS - 1), datetime.time(), datetime.timezone.utc)
    has_url = {"propertyName": "hs_linkedin_url", "operator": "HAS_PROPERTY"}
    searches = [[has_url, {"propertyName": "linkedin_checked_on", "operator": "NOT_HAS_PROPERTY"}],
                [has_url, {"propertyName": "linkedin_checked_on", "operator": "LT", "value": str(int(cutoff.timestamp() * 1000))}]]
    contacts = []
    for filters in searches:
        after = None
        while len(contacts) < MAX_PER_RUN:
            body = {"filterGroups": [{"filters": filters}], "properties": READ, "limit": 200,
                    "sorts": [{"propertyName": "linkedin_checked_on", "direction": "ASCENDING"}]}
            if after:
                body["after"] = after
            status, page = hubspot("POST", f"/crm/objects/{VERSION}/contacts/search", body)
            if status != 200:
                sys.exit(f"HubSpot {status} searching contacts: {page.get('message')}")
            contacts += page["results"]
            after = page.get("paging", {}).get("next", {}).get("after")
            if not after:
                break
            time.sleep(0.25)  # HubSpot's search takes 5 requests a second
    return contacts[:MAX_PER_RUN]


def text(value):
    return (value or "").strip()  # HarvestAPI's names and titles can end with a space


def place(city, state, country):
    return {"city": text(city), "state": text(state), "country": text(country)}


def call(provider, url):
    """One lookup. A provider failing (502, 503, 504) or Up2Data's own rate limit isn't charged: wait, then try twice more."""
    for wait in (0, 5, 15):
        time.sleep(wait)
        headers = {"Authorization": f"Token {KEY}", "X-Data-Provider": provider}
        try:
            if provider == "up2data":
                answer = requests.post(f"{API}/v1/profiles/enrich", headers=headers, json={"url": url}, timeout=90)
            else:
                answer = requests.get(f"{API}/linkedin/profile", headers=headers, params={"url": url}, timeout=90)
        except requests.RequestException as error:
            return 0, {"error": str(error)}  # no answer: the next run tries again
        try:
            body = answer.json()
        except ValueError:
            body = {}  # a gateway's error page: the status says enough
        rate_limited = answer.status_code == 429 and isinstance(body.get("error"), dict)  # Datacircle's daily limit is a string
        if not rate_limited and answer.status_code not in (502, 503, 504):
            break
    return answer.status_code, body


def check(url, provider):
    """One profile: (current jobs [(company, title)], location {city, state, country}, cost), "limit", "no profile", or None to try next run."""
    status, body = call(provider, url)
    if status == 401:
        sys.exit("401: the Datacircle API key is wrong. It's on your dashboard.")
    if status == 402:
        sys.exit("402: your balance can't cover the call. Add funds on your dashboard, then run this again.")
    if status == 429 and isinstance(body.get("error"), str):
        return "limit"
    if status == 400 or (provider == "up2data" and status == 422):
        return "no profile"  # not a profile URL, or a private or deleted profile: free
    if provider == "up2data" and status == 200:
        profile, cost = body["data"], body["datacircle_meta"]["cost_usd"]
        main, where = profile.get("current_company") or {}, profile.get("location") or {}
        jobs = [(text(main.get("name")), text(main.get("title")))] if main else []
        jobs += [(text(job.get("company")), text(job.get("title"))) for job in profile.get("positions") or [] if not job.get("ended_at")]
        return jobs, place(where.get("city"), where.get("region"), where.get("country")), cost
    if provider == "harvestapi" and status == 200:
        person, cost = body.get("element"), body["datacircle_meta"]["cost_usd"]  # null when HarvestAPI can't find the profile, still billed
        if not person:
            return [], None, cost
        where = (person.get("location") or {}).get("parsed") or {}
        held = [job for job in person.get("experience") or [] if (job.get("endDate") or {}).get("text") == "Present"]
        jobs = [(text(job.get("companyName")), text(job.get("position"))) for job in (person.get("currentPosition") or []) + held]
        return jobs, place(where.get("city"), where.get("state"), where.get("country")), cost
    reason = body.get("error", {}).get("message") if isinstance(body.get("error"), dict) else body.get("error", "")
    print(f"{url}: {status or 'no answer'} {reason}, not charged; the next run tries again", file=sys.stderr)
    return None


def plain(company):
    """A company's name to compare: "Acme, Inc." and "ACME" are the same company."""
    words = re.sub(r"[^a-z0-9]+", " ", company.lower().replace("&", " and ")).split()
    while words and words[-1] in SUFFIXES:
        words.pop()
    return " ".join(words)


def changes(contact, jobs, location):
    """The properties to write: LinkedIn's company, title and location, and the job change when the company in HubSpot is gone."""
    old = contact["properties"]
    company = (old.get("company") or "").strip()
    props = {"linkedin_checked_on": TODAY.isoformat()}
    if location is None:  # no profile found: only the date of the check
        return props, False
    kept = next((job for job in jobs if company and plain(job[0]) == plain(company)), None)  # the job HubSpot already knows, if it's still held
    props["company"], props["jobtitle"] = kept or (jobs[0] if jobs else ("", ""))
    props.update({name: value for name, value in location.items() if value})  # a place LinkedIn doesn't give leaves yours
    for name in {"jobtitle", "city", "state", "country"} - set(OVERWRITE):
        props.pop(name, None)
    moved = bool(company and not kept) or bool(old.get("linkedin_checked_on") and not company and jobs)  # a new company, gone, or a job after none
    if moved:
        props["linkedin_job_change_on"], props["linkedin_previous_company"] = TODAY.isoformat(), company
    return props, moved


def save(updates):
    for start in range(0, len(updates), 100):  # HubSpot updates 100 contacts a call
        status, body = hubspot("POST", f"/crm/objects/{VERSION}/contacts/batch/update", {"inputs": updates[start:start + 100]})
        if status == 207:
            for error in body.get("errors", []):
                print(f"HubSpot: {error.get('message')}", file=sys.stderr)
        elif status != 200:
            print(f"HubSpot {status}: {body.get('message')}; these contacts are checked again next run", file=sys.stderr)
    updates.clear()


def main(harvestapi):
    add_properties()
    contacts = due_contacts()
    provider, found, checked, moved, spent, pending = "up2data", {}, 0, 0, 0.0, []
    try:
        for contact in contacts:
            url = contact["properties"]["hs_linkedin_url"].strip()
            if url not in found:  # the same profile on two contacts is looked up once
                result = check(url, provider)
                if result == "limit" and harvestapi:
                    provider = "harvestapi"
                    result = check(url, provider)
                if result == "limit":
                    print("Up2Data's daily limit: the next run, after 00:00 UTC, goes on (or add --harvestapi to go on now through HarvestAPI).")
                    break
                if result is None:
                    continue
                found[url] = ([], None, 0) if result == "no profile" else result
                spent += found[url][2]
            jobs, location, _ = found[url]
            props, change = changes(contact, jobs, location)
            pending.append({"id": contact["id"], "properties": props})
            checked += 1
            moved += change
            if len(pending) == 100:
                save(pending)
    finally:
        save(pending)  # a stop keeps every check made
    print(f"{checked} of {len(contacts)} due contacts checked, {moved} job changes (LinkedIn job change on: {TODAY}), ${spent:.5f} spent")


if __name__ == "__main__":
    main("--harvestapi" in sys.argv[1:])

What each answer does to the run:

What the script writes or does for each answer of Datacircle's API and HubSpot's
AnswerWritten asCost
200job title, company and location; a job change when there is one$0.00125
422: private or deleted; 400: not a profile URL (a company page, say)only LinkedIn checked on, so it's tried again in 7 daysfree
429 from Up2Data's own rate limit; 502, 503, 504tried again after 5 s, then 15 s; still failing, printed and left for the next runfree
429: Datacircle's daily Up2Data limitthe run stops, and the next day's run goes on; with --harvestapi, the contact and the rest due go to HarvestAPIfree
401: wrong key; 402: balance too lowthe run stops; every check made before it is written to HubSpotfree
HubSpot 401 or 403: wrong token or a missing scopethe run stops before any check, with HubSpot's messagefree
HubSpot 429 or 5xxtried again after 10 s, then 30 sfree
HubSpot 207: a contact deleted while the script ranHubSpot's message printed; the other contacts are writtenas checked

The same LinkedIn URL on two contacts is looked up once. One run takes at most 10,000 contacts, the most HubSpot's search gives for one query; the next run takes the rest.

Run it every day

Once a day is enough, at any hour: Up2Data's daily limit starts again at 00:00 UTC, and a contact HubSpot has just updated can take a few moments to reach its search. On a Mac or Linux, crontab -e and one line, here at 7:00 every morning:

0 7 * * * cd /path/to/folder && HUBSPOT_TOKEN=... DATACIRCLE_API_KEY=... python3 enrich-hubspot-contacts-linkedin.py >> run.log 2>&1

A contact whose LinkedIn URL you fill in today is checked by the next run, before the others. To check every 30 days instead of every 7, change EVERY_DAYS at the top of the file. To keep your own job titles or locations, remove them from OVERWRITE, next to it; the company is always written, since it's what the next check compares. The script makes a handful of HubSpot calls a run: a search per 200 contacts and a save per 100, far inside a private app's limits (100 calls every 10 seconds and 250,000 a day on Free and Starter).

What it costs

Every check is a live call, at the provider's price per profile found, taken from your balance with nothing added: $0.00125 through Up2Data, $0.0037 through HarvestAPI, and $0.0023 for a HarvestAPI lookup that finds nothing. HubSpot's API costs nothing more. Up2Data takes $1 a day per account (800 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC; HarvestAPI has no daily limit. So, every profile found:

What keeping HubSpot contacts up to date from LinkedIn costs through Datacircle, every profile found
Contacts with a LinkedIn URLEach checkedThroughCost
1,000every 7 daysUp2Data$1.25 a week
5,600every 7 days (800 a day)Up2Data$7.00 a week
10,000about every 13 days (800 a day)Up2Data$7.00 a week
10,000every 7 days, with --harvestapi800 through Up2Data, 9,200 through HarvestAPI$35.04 a week

Past 5,600 contacts, the script without the flag never checks more than Up2Data's 800 a day: each contact is simply checked less often, the oldest check first. With --harvestapi, every contact due is checked that day, those past the limit through HarvestAPI. Live. Each request goes to the provider and gets the profile as it is today.

HubSpot's own data enrichment

HubSpot enriches contacts and companies itself, from its knowledge base, Get started with data enrichment (updated July 10, 2026):

HubSpot's data enrichment and this script, October 10, 2026
HubSpot data enrichmentThis script, through Datacircle
Who gets ita paid hub, Starter or above (Smart CRM and Revenue Hub: Professional or above)any HubSpot account, the free CRM included
A contact is enriched fromfirst name, last name and work emailthe contact's LinkedIn URL, its profile looked up live
A value you already havekept, unless you choose to overwrite itreplaced by LinkedIn's (the old one stays in the property's history)
How often"monthly when new data becomes available", once you turn it onevery 7 days, or what you set
Job changesnot mentionedthe date and the old company, in two properties
Pricecomes with the subscription; uses no HubSpot Credits$1.25 per 1,000 profiles found through Up2Data, $3.70 through HarvestAPI

What HubSpot's enrichment does that the script doesn't:

  • Enrich companies, from their domain, with industry and company size.
  • Enrich a contact without a LinkedIn URL, from their name and work email.
  • Run inside HubSpot, with no code, on new records as they come in.

The script gives you each contact as their LinkedIn profile shows them today, and who moved; the rest is HubSpot's.

How we tested it

We ran the file above as it is, its two hosts aside, on Python 3.9.6 and 3.12.14, against two stand-ins: one of HubSpot's API, built from its reference for version 2026-09 (the search with its filters, its sort, its pages of 200, its 10,000 results and 5 searches a second, saves of 100, a 207 for a contact deleted meanwhile, and a contact's update reaching the search 2 seconds late); one of ours, answering as its reference does, with people we made up and their jobs changed between runs. 19 contacts: a company spelled with "Inc." or "LLC", a side job, a URL with no https://, the same URL on two contacts, a private profile, a company page, a contact with no URL, one checked 3 days before, one deleted while the script ran.

  • Week 1, with a daily limit of 8 found profiles: the first run checked 9 contacts and stopped at the limit; a second run that day checked none; the next day's checked 7 more, one after Up2Data's own rate limit and one after two 503s, and left one that answered 502 three times for the next run. Three job changes: two moves and a departure (company and title emptied). None for the "Inc." and "LLC" spellings or the side job; the shared URL was looked up once; HubSpot's 207 for the deleted contact was printed, the rest written.
  • Week 2, with a limit of 3 and --harvestapi: all 17 contacts due checked, $0.04395 spent, exactly 3 × $0.00125 + 9 × $0.0037 + 3 × $0.0023 (the company page, a 400, is free). Three job changes: a move, a side job left (the company is now the main job), a new job after none. A promotion wrote the new title, the old one in the history. HubSpot answered the first save with a 429; it went through 10 seconds later.
  • Week 3: the balance ran out after 4 paid checks; the run stopped on the 402 with those checks written to HubSpot.
  • 10,250 contacts: the first run took 10,000 (50 searches, never past HubSpot's 5 a second, and 100 saves), the next run the other 250.

Both Pythons left the same contacts in the stand-in after each week. The first run on an account added the three properties; a token without the schema scopes stopped on HubSpot's 403 before anything was written. The same file, unchanged, against api.hubapi.com with a made-up token: HubSpot's 401, on paths that answer 404 with a made-up API version in them; against api.datacircle.dev with a wrong key: 401: the Datacircle API key is wrong, nothing written. We didn't run it on a real HubSpot account, having none to use: the test proves the script's logic and its calls, against the answers both APIs document.

The same weekly check for a CSV instead of a CRM: track job changes of your contacts. A whole CSV once: get LinkedIn profile data with Python; in a Google Sheet: LinkedIn profiles in Google Sheets. Both providers side by side: LinkedIn Profile API.

Questions

How do I enrich HubSpot contacts with LinkedIn data?

Read each contact's LinkedIn URL from HubSpot's API, look the profile up, and write the values back. With Datacircle: POST {"url": "<the contact's LinkedIn URL>"} to https://api.datacircle.dev/v1/profiles/enrich with your key and X-Data-Provider: up2data, then send data.current_company.title, data.current_company.name and data.location to HubSpot's batch update as jobtitle, company, city, state and country. The script on this page does it every day, each contact once a week, and flags who changed jobs.

Does HubSpot have a LinkedIn URL property?

Yes. LinkedIn URL is one of HubSpot's default contact properties, "the URL of the contact's LinkedIn page"; its internal name, the one the API uses, is hs_linkedin_url.

How much does it cost to keep 1,000 HubSpot contacts up to date every week?

$1.25 a week through Up2Data, for the profiles it finds; a private or deleted profile is free. Up2Data takes at most 800 a day per account, so up to 5,600 contacts can each be checked every week, for $7.00 a week. The $5 signup credit pays for the first 4,000 checks.

Does it work on HubSpot's free CRM?

Yes: HubSpot's contacts search and batch update APIs are on every tier, Free included, and a private app on Free or Starter gets 100 calls every 10 seconds and 250,000 a day. Creating the private app takes a super admin of the HubSpot account. HubSpot's own data enrichment needs a paid hub, Starter or above.

Is each request live, or cached?

Live. Each request goes to the provider and gets the profile as it is today.

Is there a daily limit?

Up2Data takes $1 a day per account (800 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC; HarvestAPI has no daily limit.

Sign up at datacircle.dev with your work email: a $5 credit, that's 4,000 LinkedIn profiles at $1.25 per 1,000. Free: 10M+ U.S. B2B leads, as a flat file. Download it at datacircle.dev.

Sign up