LinkedIn profiles in LlamaIndex: a LinkedIn tool through MCP or a FunctionTool
You're building an agent with LlamaIndex in Python, and it should read a LinkedIn profile from its URL: the job title, company, location and headline. Say it answers a question about someone's current role, or keeps the job title in a record up to date. LlamaIndex has no LinkedIn tool of its own. We searched its GitHub repository on October 11, 2026: among its integrations, only Bright Data's tool spec reads a LinkedIn profile, through Bright Data's own API and account.
Datacircle is a data co-op. Step 1: Query your favorite B2B data APIs through us. Same request, same price, no markup. Step 2: You're DONE. Every morning, you get the flat file of your data plus everyone else's. Right now we have 3 live LinkedIn profile APIs that we trust: Up2Data, HarvestAPI and Fetchin.
You can give your agent a LinkedIn profile tool in two ways. You can connect your agent to our MCP server and write no tool code. Or you can write one Python function that calls our API, and decide what the model reads. Both call Up2Data first, unless your agent names another provider through MCP. Up2Data costs $2.375 per 1,000 profiles it finds, and nothing for a profile it can't find. Fetchin is cheaper at $1.485 per 1,000, but it bills a profile it can't find, and all our customers share its rate limit. The function sends a URL to Fetchin only when Up2Data returns a 429, at Up2Data's daily limit or rate limit.
We ran both the MCP code and the function inside LlamaIndex's agent, FunctionAgent, with a scripted mock model and a stand-in server that answers like our API. We called api.datacircle.dev with a wrong key, from the function and from the MCP code: each got a 401, at no charge. We haven't run either with a real key or a real model.
Before you start
- Python 3.10 or later.
- LlamaIndex's core, its OpenAI models and its MCP tools:
pip install llama-index-core llama-index-llms-openai llama-index-tools-mcp. The core installs requests, which the function uses. - A Datacircle API key: Log in at datacircle.dev/login with your work email. Your API key is on the page once you're in. Put it in
DATACIRCLE_. You get a $5 credit, enough for 2,105 profiles through Up2Data.API_KEY - An OpenAI API key in
OPENAI_. The code sets the model toAPI_KEY gpt. Without that setting, LlamaIndex uses-5.6 -luna gpt, its docs say. Any model LlamaIndex can use with tools works in-3.5 -turbo llm=.
The MCP way: our MCP server through BasicMCPClient
LlamaIndex reads a remote MCP server's tools with BasicMCPClient and McpToolSpec, from the llama package, as its MCP page shows. Our MCP server is at https://. It gets a LinkedIn profile from its URL, through Up2Data, HarvestAPI or Fetchin. Our server reads your key from Authorization, so the key goes in headers as a Bearer token. This file is the whole agent:
import asyncio
import os
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
async def main():
client = BasicMCPClient(
"https://api.datacircle.dev/mcp",
headers={"Authorization": f"Bearer {os.environ['DATACIRCLE_API_KEY']}"},
timeout=60,
)
tools = await McpToolSpec(client=client, allowed_tools=["get_linkedin_profile"]).to_tool_list_async()
agent = FunctionAgent(
tools=tools,
llm=OpenAI(model="gpt-5.6-luna"),
system_prompt="Answer questions about the current role on a LinkedIn profile. Get the profile with get_linkedin_profile.",
)
response = await agent.run("What is the current job title on https://www.linkedin.com/in/williamhgates?")
print(response)
asyncio.run(main())headers and timeout aren't on LlamaIndex's MCP page, but the client's code takes both. timeout is how long the client waits for a tool's answer: 30 seconds by default, and our API waits up to 45 seconds for the provider. With the default, our test server took 35 seconds to answer, and the model got "unhandled errors in a TaskGroup (1 sub-exception)" instead of the profile. The server has other tools, such as get_, and allowed_ keeps get_ alone.
The tool takes url, and provider: up2data (the default), harvestapi or fetchin. It returns the provider's whole JSON, every job and school included, plus datacircle_meta: what the call cost and your balance after it. LlamaIndex hands the model that answer as Python prints it, so the model reads the profile twice, as JSON in text and as a dict in structured_:
meta=None content=[TextContent(type='text', text='{"data": {…}, "datacircle_meta": {…}}', …)] structured_content={'data': {…}, 'datacircle_meta': {…}} is_error=False …You pay for those tokens twice. To have the model read four short fields, use the function below. On a failed call, the model gets our API's error JSON with is_, and the run goes on. At Up2Data's limit, the server tells your agent to call again through Fetchin or HarvestAPI. The model picks one, and HarvestAPI costs more per 1,000 profiles than the other two. Our MCP server docs list every tool.
With a wrong key, the file stops at to_, before LlamaIndex calls the model, with MCPError: invalid API key or access token inside an ExceptionGroup.
The function way: one function that calls our API
You send the provider's own request to api.datacircle.dev, with your Datacircle key. That's the only change. The function sends Up2Data's own request, with X-Data-Provider naming the provider. Save the code below as datacircle_:
import json
import os
import time
import requests
API = os.environ.get("DATACIRCLE_API_URL", "https://api.datacircle.dev")
KEY = os.environ["DATACIRCLE_API_KEY"]
def call(provider, method, path, **request):
"""One call to Datacircle's API through one provider. Fetchin's 429 is its rate limit: wait a second and send it again."""
for wait in (0, 1, 2):
time.sleep(wait)
answer = requests.request(method, f"{API}{path}", headers={"Authorization": f"Token {KEY}", "X-Data-Provider": provider}, timeout=60, **request)
if provider != "fetchin" or answer.status_code != 429:
return answer
return answer
def get_linkedin_profile(url: str) -> str:
"""Get the current job title, company, location and headline on a LinkedIn profile, from the profile's URL,
like https://www.linkedin.com/in/williamhgates. Each call asks the provider live and is billed to the Datacircle balance."""
answer = call("up2data", "POST", "/v1/profiles/enrich", json={"url": url})
if answer.status_code == 200:
profile = answer.json()["data"]
company = profile.get("current_company") or {}
return json.dumps({"job_title": company.get("title"), "company": company.get("name"),
"location": (profile.get("location") or {}).get("raw"), "headline": profile.get("headline")})
if answer.status_code == 429: # Up2Data's daily limit: Fetchin answers instead
answer = call("fetchin", "GET", "/api/v1/profile", params={"profileUrlOrUrn": url})
if answer.status_code == 200:
profile = answer.json()
return json.dumps({"job_title": profile.get("jobTitle"), "company": profile.get("companyName"),
"location": profile.get("location"), "headline": profile.get("title")})
if answer.status_code in (404, 422):
return "This LinkedIn profile is private or deleted."
if answer.status_code == 400:
return "This is not a LinkedIn profile URL. Send one like https://www.linkedin.com/in/williamhgates"
if answer.status_code == 402:
return "The Datacircle balance is too low for this call. Tell the user to add funds on their Datacircle dashboard."
if answer.status_code in (429, 500, 502, 503, 504):
return f"The provider didn't answer ({answer.status_code}), and the call wasn't charged. Try again in a minute."
answer.raise_for_status() # 401: DATACIRCLE_API_KEY is wrongA FunctionAgent takes a plain Python function in tools and makes it a FunctionTool: the function's name is the tool's, its docstring is the description the model reads, and the type hint on url is its input. LlamaIndex runs a function that isn't async in a thread, so its waits don't hold up the agent. The function sends the URL to Up2Data. At Up2Data's daily limit (a 429), the function sends the same URL to Fetchin. Fetchin takes 5 requests a second across all our customers, and returns a 429 past that: the function waits one second and sends it again, then waits two seconds and sends it once more.
The model gets four fields back, as JSON text:
{"job_title": …, "company": …, "location": …, "headline": …}The function returns JSON text. A dict would reach the model through Python's str(): in our test, the model saw single quotes, which isn't JSON. For a private or deleted profile, a URL that isn't a profile, a low balance or a provider error, the tool returns a sentence, so the agent can tell the user and go on.
A wrong key (401) doesn't stop the run. If a tool raises an error, LlamaIndex gives the model the error's text as the tool's answer, here 401 Client Error: Unauthorized for url: …, and the model writes its reply from it. Set the right key in DATACIRCLE_. Each field comes from the same JSON path as in our Python post:
| Field | Up2Data's answer | Fetchin's answer |
|---|---|---|
job_ | data. | jobTitle |
company | data. | companyName |
location | data. | location |
headline | data. | title |
DATACIRCLE_ is for tests: point it at a mock of our API, and you can run the tool without spending your balance. We tested the tool against a mock.
The agent
A FunctionAgent gets its tools, a model and a system prompt, and runs the loop: the model asks for a tool, reads the tool's answer, and replies. With get_ from datacircle_:
import asyncio
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI
from datacircle_tool import get_linkedin_profile
agent = FunctionAgent(
tools=[get_linkedin_profile],
llm=OpenAI(model="gpt-5.6-luna"),
system_prompt="Answer questions about the current role on a LinkedIn profile. Get the profile with get_linkedin_profile.",
)
async def main():
response = await agent.run("What is the current job title on https://www.linkedin.com/in/williamhgates?")
print(response)
asyncio.run(main())agent.run returns a handler you await, and print(response) shows the model's reply. In the MCP example above, the agent is the same, with the MCP server's tools in tools.
Each API answer: its cost and what the tool returns
| Answer | What it means | Cost | The tool returns |
|---|---|---|---|
| Up2Data 200 | the profile | $2.375 per 1,000 | the four fields |
| Up2Data 422 | the profile is private or deleted | free | "This LinkedIn profile is private or deleted." |
| Up2Data 400 | not a LinkedIn profile URL | free | "This is not a LinkedIn profile URL." |
| Up2Data 429 | Up2Data's daily limit or rate limit | free | Fetchin's answer |
| Fetchin 200 | the profile | $1.485 per 1,000 | the four fields |
Fetchin 404 with PROFILE_ | the profile is private or deleted | $1.485 per 1,000: Fetchin bills the lookup | "This LinkedIn profile is private or deleted." |
| Fetchin 429 | its rate limit, which all our customers share | free | Fetchin's answer to a second or third try, or "Try again in a minute" |
| 402 | your balance can't cover the call | free | "The Datacircle balance is too low for this call." |
| 502, 503 or 504 | the provider failed, or didn't answer within 45 seconds | free | "Try again in a minute" |
| 401 | your key is wrong | free | the error's text, and the run goes on |
Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.
LlamaIndex's Bright Data tool
LlamaIndex's llama package has web_data_feed(source_type="linkedin_person_profile", url=...). It asks Bright Data's API for the profile, then asks again each second until the profile is ready. Bright Data bills your Bright Data account, $1.50 per 1,000 records pay as you go. Bright Data's docs say that since November 13, 2025, its LinkedIn Scraper API serves a profile's position, experience and education from a cache. Through Datacircle, each request goes to the provider and gets the profile as it is today. More differences: Bright Data LinkedIn scraper alternative.
The tests we ran
- We ran these tests on October 11, 2026, on Python 3.12, with
llama0.14.25,-index -core llama0.6.0 and-index -tools -mcp llama0.8.2, and the-index -llms -openai mcp2.3.0 and requests 2.34.2 they install. - We ran each Python file above through a real
FunctionAgent, its text unchanged. Our test code swapped OpenAI's model for LlamaIndex's mock model, which we scripted to ask forget_once, then answer with what the tool returned. It also pointed the MCP agent's URL at a stand-in server. Only the timeout test below changed a file: we took outlinkedin_ profile timeout=60. - The agent and its function ran against a stand-in server on our machine that answers like our API: the example 200 answers for Up2Data and Fetchin from our API reference, then each error in the table. Each time, the model got what the table lists for that answer.
- The MCP agent ran against a stand-in that answers like our MCP server: the client listed
get_and called it, and our errors reached the model without stopping the run. Withlinkedin_ profile timeoutleft at 30, an answer that took 35 seconds reached the model as an error. - We called api.datacircle.dev with a wrong key. Our API answered the function with a 401 and
{"error": "invalid api key"}, and the model got the error's text. The MCP agent stopped atMCPError: invalid API key or access token. These calls cost nothing. - We didn't run the Bright Data tool, or any test with a real Datacircle key or a real model.
Cost per 1,000 profiles
We charge your balance the prices on our pricing page, with no markup:
| Up2Data | HarvestAPI | Fetchin | |
|---|---|---|---|
| Per 1,000 found | $2.375 | $3.70 | $1.485 |
| Per 1,000 not found | free | $2.30 | $1.485 |
| Daily limit | 421 profiles per account | none | none |
Say your agent looks up 1,000 profiles in a day with the function, one call each, and the providers find every one. You pay us at most $1.86: $1.00 for 421 through Up2Data and $0.86 for the other 579 through Fetchin. An agent may call the tool more than once for a question, and we bill each call. OpenAI bills you for the model's tokens.
The same tool for a LangChain agent: LinkedIn profiles in LangChain. For a CrewAI crew: LinkedIn profiles in CrewAI. For OpenAI's Agents SDK: LinkedIn profiles in the OpenAI Agents SDK. For a Pydantic AI agent: LinkedIn profiles in Pydantic AI. The same call from a Python script: get LinkedIn profile data with Python. In Claude, ChatGPT or Cursor: our LinkedIn MCP server. Other vendors' prices per 1,000: LinkedIn profile API pricing compared.
Questions
Does LlamaIndex have a LinkedIn tool?
Not one of its own. Among its integrations, Bright Data's tool spec reads a LinkedIn profile through Bright Data's API, and Bright Data bills your account. To use ours, add our MCP server at https://
How do I connect a LlamaIndex agent to an MCP server with an API key?
Create BasicMCPClient("https://
How much does a LinkedIn profile cost?
$2.375 per 1,000 through Up2Data (a profile it can't find is free), $3.70 per 1,000 through HarvestAPI. $1.485 per 1,000 through Fetchin, a profile it can't find billed the same. HarvestAPI bills a profile it can't find at $2.30 per 1,000.
Is there a daily limit?
Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.
Is each request live, or cached?
Live. Each request goes to the provider and gets the profile as it is today.
Do I need a LinkedIn account?
No. You send the profile's URL with your Datacircle key: no LinkedIn login, no cookies, no browser.
What happens when my balance runs out?
A call your balance can't cover answers 402. Add funds, from $5, on your dashboard.
Get started
Sign up at datacircle.dev with your work email: a $5 credit, that's 2,105 LinkedIn profiles at $2.375 per 1,000.
Sign up