LinkedIn profiles in Haystack: a LinkedIn tool through MCP or a Python function
You're building an agent in Python with Haystack, and the model should read a LinkedIn profile from its URL: the job title, company, location and headline. Haystack's README calls it "an open-source AI orchestration framework for building production-ready LLM applications in Python". It has no LinkedIn tool of its own. We searched the GitHub code of deepset, the company behind Haystack, for "linkedin" on October 11, 2026: we found links to authors' profiles and test files, and no tool.
Datacircle is a data co-op. Step 1: Query your favorite B2B data APIs through us. Same request, same price, no markup. Step 2: You're DONE. Every morning, you get the flat file of your data plus everyone else's. Add $50 to your account: you get $50 of API PLUS the flat file. Right now we have 3 live LinkedIn profile APIs that we trust: Up2Data, HarvestAPI and Fetchin.
You can connect our MCP server to your agent and write no tool code. Or you can write one Python function that calls our API, and decide what the model reads. Both call Up2Data first. Up2Data costs $2.375 per 1,000 profiles it finds, and nothing for a profile it can't find. Fetchin is cheaper at $1.485 per 1,000, but it charges for a profile it can't find, and all our customers share its rate limit. Both switch to Fetchin only when Up2Data hits its daily limit and returns a 429.
We ran both agents inside Haystack, with a scripted stand-in for OpenAI's API as the model, and a stand-in server that answers like our API. We called api.datacircle.dev with a wrong key, from the function and from the MCP code: each got a 401, at no charge. We haven't run either with a real key or a real model.
LinkedIn in Haystack's tools and integrations
A Haystack Agent calls Python functions through @tool, Haystack components, pipelines, other agents, and an MCP server's tools through the mcp-haystack package. None of them is a LinkedIn tool, and Haystack's integrations page lists none for LinkedIn by name.
One integration reads LinkedIn: Bright Data's. It calls Bright Data's API with your Bright Data key, and Bright Data bills you. Our comparison: Bright Data LinkedIn scraper alternative.
Before you start
- Python 3.10 or later. We tested on 3.12.
- Haystack, its MCP integration if you connect over MCP, and httpx if you call our API from a function:
pip install haystack-ai mcp-haystack httpx - A Datacircle API key: Log in at datacircle.dev/login with your work email. Your API key is on the page once you're in. Put it in
DATACIRCLE_. You get a $5 credit, enough for 2,105 profiles through Up2Data.API_KEY - An OpenAI API key in
OPENAI_, for Haystack'sAPI_KEY OpenAIChatGenerator. Haystack's Tool page says its tools give "a consistent experience across models": with another of its chat generators, the tools below stay the same, and only thechat_generatorline and its import change. We haven't tried another one.
The MCP way: our MCP server through MCPToolset
mcp-haystack connects to an MCP server with MCPToolset, as Haystack's MCPToolset page shows. Given a token, it sends your key to our server as a Bearer token, with each request. Our MCP server is at https://. It gets a LinkedIn profile from its URL, through Up2Data, HarvestAPI or Fetchin. Save this file as mcp_ and run python mcp_agent.py:
import json
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.utils import Secret
from haystack_integrations.tools.mcp import MCPToolset, StreamableHttpServerInfo
def our_text(result):
"""Our server's answer, once: without this, the model reads the profile twice."""
return json.loads(result)["content"][0]["text"]
datacircle = MCPToolset(
server_info=StreamableHttpServerInfo(
url="https://api.datacircle.dev/mcp",
token=Secret.from_env_var("DATACIRCLE_API_KEY"),
timeout=60, # the HTTP client's timeout: above the 45 seconds our API waits for a provider
),
tool_names=["get_linkedin_profile"],
invocation_timeout=60, # the tool call's own timeout, also above 45 seconds
outputs_to_string={"get_linkedin_profile": {"handler": our_text}},
)
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-5.4-mini"),
tools=datacircle,
system_prompt="Get a profile with get_linkedin_profile. If up2data says its limit is reached, call it again with provider fetchin.",
max_agent_steps=4, # at most 4 rounds with the model: each tool call is a call to our API
)
result = agent.run(messages=[ChatMessage.from_user("What is the current job title on https://www.linkedin.com/in/example-profile?")])
print(result["last_message"].text if result["exit_reason"] == "text" else f"Stopped: {result['exit_reason']}")
datacircle.close()MCPToolset(... doesn't connect. agent.run does, before it calls the model: it opens the connection and lists the tools. datacircle.close() closes it.
The server has other tools, such as get_, and tool_names keeps get_ alone. With a name our server doesn't have, agent.run raised MCPToolNotFoundError and listed the server's tool names, before any model call. The first request to the model was 1,816 bytes with the one tool, and 3,930 bytes with all five.
mcp-haystack's MCPTool, for one tool, sent the model an empty description: it connects on first use and keeps the empty description it had before it connected, unless you pass eager_. Our description tells the model when to switch to Fetchin or HarvestAPI, so this file uses MCPToolset.
Without a handler, the model reads each profile twice
Our server sends the answer twice in one result: as text, and as structuredContent. By default, Haystack gives the model both:
{"_meta":null,"content":[{"type":"text","text":"{\"data\": …}"}],"structuredContent":{"data":{…},…},"isError":false,"resultType":"complete"}Through Up2Data, our answer is 1,167 characters, and the model read 2,445. A saved Fetchin answer of 65,229 characters reached the model as 132,169. With the our_text handler in outputs_to_string, the model got each answer once and whole: the provider's JSON, every job and school included, plus datacircle_meta, what the call cost and your balance after it:
{"data": {…}, "meta": {…}, "datacircle_meta": {…}}Haystack sends the model each tool's name, description and inputs, and leaves out its output schema. Our server describes the answer in a schema of 8,291 characters, and none of it reached the model.
At Up2Data's limit, the model reads our error inside a Python error
The tool takes url, and provider: up2data (the default), harvestapi or fetchin. At Up2Data's limit, the server tells your agent to call again through Fetchin or HarvestAPI. HarvestAPI costs more per 1,000 profiles than the other two, so we name Fetchin in the file's prompt. Our server marks Up2Data's 429 as an error, and Haystack turns an error into this text for the model:
Failed to invoke Tool `get_linkedin_profile` with parameters {…}. Error: Failed to invoke tool 'get_linkedin_profile' with args: {…}, got error: Tool 'get_linkedin_profile' returned an error: [TextContent(type='text', text='{"error": …}', annotations=None, meta=None)]That text holds our error JSON but no status code. When Up2Data hit its limit, our JSON said so, and our scripted model called again with fetchin and got the profile. The other errors we tried reached the model the same way, and the run went on. Our MCP server docs list every tool.
A wrong key stops agent.run before the model runs
With a wrong key, our server answered the first request with a 401, and agent.run raised before any model call. We tried it on api.datacircle.dev:
mcp.shared.exceptions.MCPError: invalid API key or access token
…
haystack_integrations.tools.mcp.mcp_tool.MCPConnectionError: Failed to connect to MCP server via streamable HTTP. Please check if:
…
3. Authentication token is correct (if required)The last error doesn't say why. Our message, invalid API key or access token, is higher up. Set the right key in DATACIRCLE_. We haven't tried an OAuth sign in through mcp-haystack.
Set both timeouts above the 45 seconds our API waits for a provider
mcp-haystack has two timeouts, each 30 seconds by default: timeout on the HTTP client in StreamableHttpServerInfo, and invocation_ on the call. Our stand-in server waited 50 seconds before it answered. With both at 60 seconds, the model got the profile. With the defaults, after 30 seconds the model read this, with nothing after Error::
Failed to invoke Tool `get_linkedin_profile` with parameters {'url': 'https://www.linkedin.com/in/example-profile', 'provider': 'up2data'}. Error: With invocation_ alone raised, the same. With timeout alone, the call stopped at 30 seconds with Error: Operation timed out after 30.0 seconds.
The function way: one Python function that calls our API
You send the provider's own request to api.datacircle.dev, with your Datacircle key. That's the only change. The function sends Up2Data's and Fetchin's own requests, with X-Data-Provider naming the provider. Up2Data's request is in our API reference. Save this file as tool_ and run python tool_agent.py:
import os
import time
from typing import Annotated
import httpx
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import tool
API = os.environ.get("DATACIRCLE_API_URL", "https://api.datacircle.dev")
KEY = os.environ["DATACIRCLE_API_KEY"]
def call(provider, method, path, **request):
"""One call to Datacircle's API through one provider. Fetchin's 429 is its rate limit: wait a second and send it again."""
for wait in (0, 1, 2):
time.sleep(wait)
answer = httpx.request(method, f"{API}{path}", headers={"Authorization": f"Token {KEY}", "X-Data-Provider": provider}, timeout=60, **request)
if provider != "fetchin" or answer.status_code != 429:
return answer
return answer
@tool
def get_linkedin_profile(url: Annotated[str, "The profile's LinkedIn URL, like https://www.linkedin.com/in/example-profile"]) -> dict:
"""Get the current job title, company, location and headline on a LinkedIn profile, from the profile's URL.
Returns the keys job_title, company, location and headline, or one key, error, when the profile can't be read.
Each call asks the provider live and may be billed to the Datacircle balance."""
answer = call("up2data", "POST", "/v1/profiles/enrich", json={"url": url})
if answer.status_code == 200:
profile = answer.json()["data"]
company = profile.get("current_company") or {}
return {"job_title": company.get("title"), "company": company.get("name"),
"location": (profile.get("location") or {}).get("raw"), "headline": profile.get("headline")}
if answer.status_code == 429: # Up2Data's daily limit: Fetchin answers instead
answer = call("fetchin", "GET", "/api/v1/profile", params={"profileUrlOrUrn": url})
if answer.status_code == 200:
profile = answer.json()
return {"job_title": profile.get("jobTitle"), "company": profile.get("companyName"),
"location": profile.get("location"), "headline": profile.get("title")}
if answer.status_code in (404, 422):
return {"error": "This LinkedIn profile is private or deleted."}
if answer.status_code == 400:
return {"error": "This is not a LinkedIn profile URL. Send one like https://www.linkedin.com/in/example-profile"}
if answer.status_code == 402:
return {"error": "The Datacircle balance is too low for this call. Tell the user to add funds on their Datacircle dashboard."}
if answer.status_code in (429, 500, 502, 503, 504):
return {"error": f"The provider didn't answer ({answer.status_code}), and the call wasn't charged. Try again in a minute."}
answer.raise_for_status() # 401: DATACIRCLE_API_KEY is wrong
agent = Agent(
chat_generator=OpenAIChatGenerator(model="gpt-5.4-mini"),
tools=[get_linkedin_profile],
max_agent_steps=4, # at most 4 rounds with the model: each tool call is a call to our API
)
result = agent.run(messages=[ChatMessage.from_user("What is the current job title on https://www.linkedin.com/in/example-profile?")])
print(result["last_message"].text if result["exit_reason"] == "text" else f"Stopped: {result['exit_reason']}")Haystack's @tool turns the function into a tool. The model gets the function's name, its whole docstring as the description, and the Annotated text as the description of url. The model never sees -> dict, so we name the keys of the answer in the docstring.
The function sends the URL to Up2Data. At Up2Data's daily limit (its 429), it sends the same URL to Fetchin. Fetchin takes 5 requests a second across all our customers, and returns a 429 past that: the function waits one second and sends it again, then waits two seconds and sends it once more.
The function returns a dict, and Haystack gives it to the model as JSON text:
{"job_title": …, "company": …, "location": …, "headline": …}For a private or deleted profile, a URL that isn't a profile, a low balance or a provider error, the function returns one key, error, so the agent can tell the user and go on. Each field comes from the same JSON path as in our Python post:
| Field | Up2Data's answer | Fetchin's answer |
|---|---|---|
job_ | data. | jobTitle |
company | data. | companyName |
location | data. | location |
headline | data. | title |
With a wrong key (401), the function raises. Haystack prints the error, gives it to the model as the tool's result, and keeps the run going. We tried a wrong key on api.datacircle.dev, and the model read this:
Failed to invoke Tool `get_linkedin_profile` with parameters {'url': 'https://www.linkedin.com/in/example-profile'}. Error: Client error '401 Unauthorized' for url 'https://api.datacircle.dev/v1/profiles/enrich'
For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/401Set the right key in DATACIRCLE_. To stop the run on an error instead, pass raise_ to the Agent. Haystack sets no time limit on a function, and httpx waits 60 seconds here: our stand-in server waited 50 seconds before it answered, and the model got the answer. DATACIRCLE_ is for tests: point it at a stand-in for our API, and you can run the function without spending your balance.
Cap the agent's calls with max_agent_steps
Each tool call is a call to our API. We bill each call that gets a profile, and each call for a profile Fetchin or HarvestAPI can't find, at the price on our pricing page. The function tool sends up to four requests per call, and we bill at most one of them: Up2Data's 429 and Fetchin's 429 are free.
Haystack's Agent stops after max_agent_steps, 100 by default. A step is one call to the model, plus every tool call it asked for. A scripted model that asked for the tool at every turn made 100 calls to a stand-in for our API in about a second. With max_agent_steps=4, as in both files, it got 4 profiles, and 8 when it asked for two at a time. Then exit_reason is max_agent_steps and the last message is the tool's result, with no answer from the model, so both files check exit_reason.
By default, Haystack calls the tool without asking you. Its human in the loop page adds a ConfirmationHook. We added this import and this argument to the Agent in mcp_:
from haystack.hooks.human_in_the_loop import AlwaysAskPolicy, BlockingConfirmationStrategy, ConfirmationHook, SimpleConsoleUI
hooks={"before_tool": [ConfirmationHook(confirmation_strategies={"get_linkedin_profile": BlockingConfirmationStrategy(confirmation_policy=AlwaysAskPolicy(), confirmation_ui=SimpleConsoleUI())})]},Before each call, the script printed this and waited for an answer:
--- Tool Execution Request ---
Tool: get_linkedin_profile
Description: …
Arguments:
url: https://www.linkedin.com/in/example-profile
provider: up2data
------------------------------
Confirm execution? (y=confirm / n=reject / m=modify):With n, nothing reached our server, and the model read "Tool execution for 'get_linkedin_profile' was rejected by the user." With no terminal to answer, the script stopped on EOFError.
API answers: cost and what the tool returns
| Answer | Meaning | Cost | The tool returns |
|---|---|---|---|
| Up2Data 200 | the profile | $2.375 per 1,000 | the four fields |
| Up2Data 422 | the profile is private or deleted | free | error: "This LinkedIn profile is private or deleted." |
| Up2Data 400 | not a LinkedIn profile URL | free | error: "This is not a LinkedIn profile URL." |
| Up2Data 429 | its daily limit | free | Fetchin's answer |
| Fetchin 200 | the profile | $1.485 per 1,000 | the four fields |
Fetchin 404 with PROFILE_ | the profile is private or deleted | $1.485 per 1,000: Fetchin bills the lookup | error: "This LinkedIn profile is private or deleted." |
| Fetchin 429 | its rate limit, which all our customers share | free | Fetchin's answer after up to two retries, or error: "Try again in a minute" |
| 402 | your balance can't cover the call | free | error: "The Datacircle balance is too low for this call." |
| 500, 502, 503 or 504 | the provider failed, or didn't answer within 45 seconds | free | error: "Try again in a minute" |
| 401 | your key is wrong | free | an exception, which the model reads as Failed to invoke Tool and the error message |
Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.
The tests we ran
- We tested on October 11, 2026, with Python 3.12,
haystack-ai3.3.0,mcp-haystack1.5.2,mcp2.3.0, openai 3.28.0 and httpx 0.28.1. - We ran each file and each variant above in a real Haystack agent, with
OPENAI_on a scripted stand-in for OpenAI's API.BASE_URL - We ran the function against a stand-in for our API that returned each answer in the table, and the MCP agent against a stand-in for our MCP server with the tools our live server lists. Through MCP, we tried each error, a 50 second answer, the long answer, a wrong tool name and the approval step.
- We didn't try another chat generator, an OAuth sign in, a real Datacircle key or a real model.
Cost per 1,000 profiles
We charge your balance the prices on our pricing page, with no markup:
| Up2Data | HarvestAPI | Fetchin | |
|---|---|---|---|
| Per 1,000 found | $2.375 | $3.70 | $1.485 |
| Per 1,000 not found | free | $2.30 | $1.485 |
| Daily limit | 421 profiles per account | none | none |
Say your agent looks up 1,000 profiles in a day, one call each, and the providers find every one. You pay us at most $1.86: $1.00 for 421 through Up2Data and $0.86 for the other 579 through Fetchin. Through Fetchin alone, the same 1,000 would cost less: $1.485. Both files call Up2Data first because it bills nothing for a profile it can't find, and all our customers share Fetchin's rate limit. Your agent may call the tool more than once per question. Your model's provider bills you for its tokens.
The same tool in other agent frameworks: LinkedIn profiles in LangChain, LinkedIn profiles in Pydantic AI, LinkedIn profiles in Google ADK, LinkedIn profiles in the Claude Agent SDK, LinkedIn profiles in smolagents, LinkedIn profiles in Agno, LinkedIn profiles in Microsoft Agent Framework, LinkedIn profiles in Strands Agents and LinkedIn profiles in DSPy. The same call from a Python script: get LinkedIn profile data with Python. In Claude, ChatGPT or Cursor: our LinkedIn MCP server. Other vendors' prices per 1,000: LinkedIn profile API pricing compared.
Questions
Does Haystack have a LinkedIn tool?
Not one of its own. Bright Data's Haystack integration reads a LinkedIn profile with your Bright Data key. To read profiles through us, connect our MCP server at https://
How do I send an API key to a remote MCP server in Haystack?
Pass it as the token of StreamableHttpServerInfo: StreamableHttpServerInfo(url="https://
Why does my Haystack MCP tool fail after 30 seconds?
mcp-haystack waits 30 seconds by default, twice: the HTTP client's timeout in StreamableHttpServerInfo(timeout=...
How much does a LinkedIn profile cost?
$2.375 per 1,000 through Up2Data (a profile it can't find is free), $3.70 per 1,000 through HarvestAPI. $1.485 per 1,000 through Fetchin, a profile it can't find billed the same. HarvestAPI bills a profile it can't find at $2.30 per 1,000.
Is there a daily limit?
Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.
Is each request live, or cached?
Live. Each request goes to the provider and gets the profile as it is today.
Do I need a LinkedIn account?
No. You send the profile's URL with your Datacircle key: no LinkedIn login, no cookies, no browser.
What happens when my balance runs out?
A call your balance can't cover answers 402. Add funds, from $5, on your dashboard.
Get started
Sign up at datacircle.dev with your work email: a $5 credit, that's 2,105 LinkedIn profiles at $2.375 per 1,000.
Sign up