Datacircle

LinkedIn profiles in Haystack: a LinkedIn tool through MCP or a Python function

You're building an agent in Python with Haystack, and the model should read a LinkedIn profile from its URL: the job title, company, location and headline. Haystack's README calls it "an open-source AI orchestration framework for building production-ready LLM applications in Python". It has no LinkedIn tool of its own. We searched the GitHub code of deepset, the company behind Haystack, for "linkedin" on October 11, 2026: we found links to authors' profiles and test files, and no tool.

Datacircle is a data co-op. Step 1: Query your favorite B2B data APIs through us. Same request, same price, no markup. Step 2: You're DONE. Every morning, you get the flat file of your data plus everyone else's. Add $50 to your account: you get $50 of API PLUS the flat file. Right now we have 3 live LinkedIn profile APIs that we trust: Up2Data, HarvestAPI and Fetchin.

You can connect our MCP server to your agent and write no tool code. Or you can write one Python function that calls our API, and decide what the model reads. Both call Up2Data first. Up2Data costs $2.375 per 1,000 profiles it finds, and nothing for a profile it can't find. Fetchin is cheaper at $1.485 per 1,000, but it charges for a profile it can't find, and all our customers share its rate limit. Both switch to Fetchin only when Up2Data hits its daily limit and returns a 429.

We ran both agents inside Haystack, with a scripted stand-in for OpenAI's API as the model, and a stand-in server that answers like our API. We called api.datacircle.dev with a wrong key, from the function and from the MCP code: each got a 401, at no charge. We haven't run either with a real key or a real model.

LinkedIn in Haystack's tools and integrations

A Haystack Agent calls Python functions through @tool, Haystack components, pipelines, other agents, and an MCP server's tools through the mcp-haystack package. None of them is a LinkedIn tool, and Haystack's integrations page lists none for LinkedIn by name.

One integration reads LinkedIn: Bright Data's. It calls Bright Data's API with your Bright Data key, and Bright Data bills you. Our comparison: Bright Data LinkedIn scraper alternative.

Before you start

  • Python 3.10 or later. We tested on 3.12.
  • Haystack, its MCP integration if you connect over MCP, and httpx if you call our API from a function:
    pip install haystack-ai mcp-haystack httpx
  • A Datacircle API key: Log in at datacircle.dev/login with your work email. Your API key is on the page once you're in. Put it in DATACIRCLE_API_KEY. You get a $5 credit, enough for 2,105 profiles through Up2Data.
  • An OpenAI API key in OPENAI_API_KEY, for Haystack's OpenAIChatGenerator. Haystack's Tool page says its tools give "a consistent experience across models": with another of its chat generators, the tools below stay the same, and only the chat_generator line and its import change. We haven't tried another one.

The MCP way: our MCP server through MCPToolset

mcp-haystack connects to an MCP server with MCPToolset, as Haystack's MCPToolset page shows. Given a token, it sends your key to our server as a Bearer token, with each request. Our MCP server is at https://api.datacircle.dev/mcp. It gets a LinkedIn profile from its URL, through Up2Data, HarvestAPI or Fetchin. Save this file as mcp_agent.py and run python mcp_agent.py:

import json

from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.utils import Secret
from haystack_integrations.tools.mcp import MCPToolset, StreamableHttpServerInfo


def our_text(result):
    """Our server's answer, once: without this, the model reads the profile twice."""
    return json.loads(result)["content"][0]["text"]


datacircle = MCPToolset(
    server_info=StreamableHttpServerInfo(
        url="https://api.datacircle.dev/mcp",
        token=Secret.from_env_var("DATACIRCLE_API_KEY"),
        timeout=60,  # the HTTP client's timeout: above the 45 seconds our API waits for a provider
    ),
    tool_names=["get_linkedin_profile"],
    invocation_timeout=60,  # the tool call's own timeout, also above 45 seconds
    outputs_to_string={"get_linkedin_profile": {"handler": our_text}},
)

agent = Agent(
    chat_generator=OpenAIChatGenerator(model="gpt-5.4-mini"),
    tools=datacircle,
    system_prompt="Get a profile with get_linkedin_profile. If up2data says its limit is reached, call it again with provider fetchin.",
    max_agent_steps=4,  # at most 4 rounds with the model: each tool call is a call to our API
)
result = agent.run(messages=[ChatMessage.from_user("What is the current job title on https://www.linkedin.com/in/example-profile?")])
print(result["last_message"].text if result["exit_reason"] == "text" else f"Stopped: {result['exit_reason']}")
datacircle.close()

MCPToolset(...) doesn't connect. agent.run does, before it calls the model: it opens the connection and lists the tools. datacircle.close() closes it.

The server has other tools, such as get_balance, and tool_names keeps get_linkedin_profile alone. With a name our server doesn't have, agent.run raised MCPToolNotFoundError and listed the server's tool names, before any model call. The first request to the model was 1,816 bytes with the one tool, and 3,930 bytes with all five.

mcp-haystack's MCPTool, for one tool, sent the model an empty description: it connects on first use and keeps the empty description it had before it connected, unless you pass eager_connect=True. Our description tells the model when to switch to Fetchin or HarvestAPI, so this file uses MCPToolset.

Without a handler, the model reads each profile twice

Our server sends the answer twice in one result: as text, and as structuredContent. By default, Haystack gives the model both:

{"_meta":null,"content":[{"type":"text","text":"{\"data\": …}"}],"structuredContent":{"data":{…},…},"isError":false,"resultType":"complete"}

Through Up2Data, our answer is 1,167 characters, and the model read 2,445. A saved Fetchin answer of 65,229 characters reached the model as 132,169. With the our_text handler in outputs_to_string, the model got each answer once and whole: the provider's JSON, every job and school included, plus datacircle_meta, what the call cost and your balance after it:

{"data": {…}, "meta": {…}, "datacircle_meta": {…}}

Haystack sends the model each tool's name, description and inputs, and leaves out its output schema. Our server describes the answer in a schema of 8,291 characters, and none of it reached the model.

At Up2Data's limit, the model reads our error inside a Python error

The tool takes url, and provider: up2data (the default), harvestapi or fetchin. At Up2Data's limit, the server tells your agent to call again through Fetchin or HarvestAPI. HarvestAPI costs more per 1,000 profiles than the other two, so we name Fetchin in the file's prompt. Our server marks Up2Data's 429 as an error, and Haystack turns an error into this text for the model:

Failed to invoke Tool `get_linkedin_profile` with parameters {…}. Error: Failed to invoke tool 'get_linkedin_profile' with args: {…}, got error: Tool 'get_linkedin_profile' returned an error: [TextContent(type='text', text='{"error": …}', annotations=None, meta=None)]

That text holds our error JSON but no status code. When Up2Data hit its limit, our JSON said so, and our scripted model called again with fetchin and got the profile. The other errors we tried reached the model the same way, and the run went on. Our MCP server docs list every tool.

A wrong key stops agent.run before the model runs

With a wrong key, our server answered the first request with a 401, and agent.run raised before any model call. We tried it on api.datacircle.dev:

mcp.shared.exceptions.MCPError: invalid API key or access token
…
haystack_integrations.tools.mcp.mcp_tool.MCPConnectionError: Failed to connect to MCP server via streamable HTTP. Please check if:
…
3. Authentication token is correct (if required)

The last error doesn't say why. Our message, invalid API key or access token, is higher up. Set the right key in DATACIRCLE_API_KEY. We haven't tried an OAuth sign in through mcp-haystack.

Set both timeouts above the 45 seconds our API waits for a provider

mcp-haystack has two timeouts, each 30 seconds by default: timeout on the HTTP client in StreamableHttpServerInfo, and invocation_timeout on the call. Our stand-in server waited 50 seconds before it answered. With both at 60 seconds, the model got the profile. With the defaults, after 30 seconds the model read this, with nothing after Error::

Failed to invoke Tool `get_linkedin_profile` with parameters {'url': 'https://www.linkedin.com/in/example-profile', 'provider': 'up2data'}. Error: 

With invocation_timeout alone raised, the same. With timeout alone, the call stopped at 30 seconds with Error: Operation timed out after 30.0 seconds.

The function way: one Python function that calls our API

You send the provider's own request to api.datacircle.dev, with your Datacircle key. That's the only change. The function sends Up2Data's and Fetchin's own requests, with X-Data-Provider naming the provider. Up2Data's request is in our API reference. Save this file as tool_agent.py and run python tool_agent.py:

import os
import time
from typing import Annotated

import httpx
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.tools import tool

API = os.environ.get("DATACIRCLE_API_URL", "https://api.datacircle.dev")
KEY = os.environ["DATACIRCLE_API_KEY"]


def call(provider, method, path, **request):
    """One call to Datacircle's API through one provider. Fetchin's 429 is its rate limit: wait a second and send it again."""
    for wait in (0, 1, 2):
        time.sleep(wait)
        answer = httpx.request(method, f"{API}{path}", headers={"Authorization": f"Token {KEY}", "X-Data-Provider": provider}, timeout=60, **request)
        if provider != "fetchin" or answer.status_code != 429:
            return answer
    return answer


@tool
def get_linkedin_profile(url: Annotated[str, "The profile's LinkedIn URL, like https://www.linkedin.com/in/example-profile"]) -> dict:
    """Get the current job title, company, location and headline on a LinkedIn profile, from the profile's URL.
    Returns the keys job_title, company, location and headline, or one key, error, when the profile can't be read.
    Each call asks the provider live and may be billed to the Datacircle balance."""
    answer = call("up2data", "POST", "/v1/profiles/enrich", json={"url": url})
    if answer.status_code == 200:
        profile = answer.json()["data"]
        company = profile.get("current_company") or {}
        return {"job_title": company.get("title"), "company": company.get("name"),
                "location": (profile.get("location") or {}).get("raw"), "headline": profile.get("headline")}
    if answer.status_code == 429:  # Up2Data's daily limit: Fetchin answers instead
        answer = call("fetchin", "GET", "/api/v1/profile", params={"profileUrlOrUrn": url})
        if answer.status_code == 200:
            profile = answer.json()
            return {"job_title": profile.get("jobTitle"), "company": profile.get("companyName"),
                    "location": profile.get("location"), "headline": profile.get("title")}
    if answer.status_code in (404, 422):
        return {"error": "This LinkedIn profile is private or deleted."}
    if answer.status_code == 400:
        return {"error": "This is not a LinkedIn profile URL. Send one like https://www.linkedin.com/in/example-profile"}
    if answer.status_code == 402:
        return {"error": "The Datacircle balance is too low for this call. Tell the user to add funds on their Datacircle dashboard."}
    if answer.status_code in (429, 500, 502, 503, 504):
        return {"error": f"The provider didn't answer ({answer.status_code}), and the call wasn't charged. Try again in a minute."}
    answer.raise_for_status()  # 401: DATACIRCLE_API_KEY is wrong


agent = Agent(
    chat_generator=OpenAIChatGenerator(model="gpt-5.4-mini"),
    tools=[get_linkedin_profile],
    max_agent_steps=4,  # at most 4 rounds with the model: each tool call is a call to our API
)
result = agent.run(messages=[ChatMessage.from_user("What is the current job title on https://www.linkedin.com/in/example-profile?")])
print(result["last_message"].text if result["exit_reason"] == "text" else f"Stopped: {result['exit_reason']}")

Haystack's @tool turns the function into a tool. The model gets the function's name, its whole docstring as the description, and the Annotated text as the description of url. The model never sees -> dict, so we name the keys of the answer in the docstring.

The function sends the URL to Up2Data. At Up2Data's daily limit (its 429), it sends the same URL to Fetchin. Fetchin takes 5 requests a second across all our customers, and returns a 429 past that: the function waits one second and sends it again, then waits two seconds and sends it once more.

The function returns a dict, and Haystack gives it to the model as JSON text:

{"job_title": …, "company": …, "location": …, "headline": …}

For a private or deleted profile, a URL that isn't a profile, a low balance or a provider error, the function returns one key, error, so the agent can tell the user and go on. Each field comes from the same JSON path as in our Python post:

Source of each field
FieldUp2Data's answerFetchin's answer
job_titledata.current_company.titlejobTitle
companydata.current_company.namecompanyName
locationdata.location.rawlocation
headlinedata.headlinetitle

With a wrong key (401), the function raises. Haystack prints the error, gives it to the model as the tool's result, and keeps the run going. We tried a wrong key on api.datacircle.dev, and the model read this:

Failed to invoke Tool `get_linkedin_profile` with parameters {'url': 'https://www.linkedin.com/in/example-profile'}. Error: Client error '401 Unauthorized' for url 'https://api.datacircle.dev/v1/profiles/enrich'
For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/401

Set the right key in DATACIRCLE_API_KEY. To stop the run on an error instead, pass raise_on_tool_invocation_failure=True to the Agent. Haystack sets no time limit on a function, and httpx waits 60 seconds here: our stand-in server waited 50 seconds before it answered, and the model got the answer. DATACIRCLE_API_URL is for tests: point it at a stand-in for our API, and you can run the function without spending your balance.

Cap the agent's calls with max_agent_steps

Each tool call is a call to our API. We bill each call that gets a profile, and each call for a profile Fetchin or HarvestAPI can't find, at the price on our pricing page. The function tool sends up to four requests per call, and we bill at most one of them: Up2Data's 429 and Fetchin's 429 are free.

Haystack's Agent stops after max_agent_steps, 100 by default. A step is one call to the model, plus every tool call it asked for. A scripted model that asked for the tool at every turn made 100 calls to a stand-in for our API in about a second. With max_agent_steps=4, as in both files, it got 4 profiles, and 8 when it asked for two at a time. Then exit_reason is max_agent_steps and the last message is the tool's result, with no answer from the model, so both files check exit_reason.

By default, Haystack calls the tool without asking you. Its human in the loop page adds a ConfirmationHook. We added this import and this argument to the Agent in mcp_agent.py:

from haystack.hooks.human_in_the_loop import AlwaysAskPolicy, BlockingConfirmationStrategy, ConfirmationHook, SimpleConsoleUI

    hooks={"before_tool": [ConfirmationHook(confirmation_strategies={"get_linkedin_profile": BlockingConfirmationStrategy(confirmation_policy=AlwaysAskPolicy(), confirmation_ui=SimpleConsoleUI())})]},

Before each call, the script printed this and waited for an answer:

--- Tool Execution Request ---
Tool: get_linkedin_profile
Description: …
Arguments:
  url: https://www.linkedin.com/in/example-profile
  provider: up2data
------------------------------
Confirm execution? (y=confirm / n=reject / m=modify):

With n, nothing reached our server, and the model read "Tool execution for 'get_linkedin_profile' was rejected by the user." With no terminal to answer, the script stopped on EOFError.

API answers: cost and what the tool returns

Our API's answers to the function in tool_agent.py
AnswerMeaningCostThe tool returns
Up2Data 200the profile$2.375 per 1,000the four fields
Up2Data 422the profile is private or deletedfreeerror: "This LinkedIn profile is private or deleted."
Up2Data 400not a LinkedIn profile URLfreeerror: "This is not a LinkedIn profile URL."
Up2Data 429its daily limitfreeFetchin's answer
Fetchin 200the profile$1.485 per 1,000the four fields
Fetchin 404 with PROFILE_NOT_FOUNDthe profile is private or deleted$1.485 per 1,000: Fetchin bills the lookuperror: "This LinkedIn profile is private or deleted."
Fetchin 429its rate limit, which all our customers sharefreeFetchin's answer after up to two retries, or error: "Try again in a minute"
402your balance can't cover the callfreeerror: "The Datacircle balance is too low for this call."
500, 502, 503 or 504the provider failed, or didn't answer within 45 secondsfreeerror: "Try again in a minute"
401your key is wrongfreean exception, which the model reads as Failed to invoke Tool and the error message

Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.

The tests we ran

  • We tested on October 11, 2026, with Python 3.12, haystack-ai 3.3.0, mcp-haystack 1.5.2, mcp 2.3.0, openai 3.28.0 and httpx 0.28.1.
  • We ran each file and each variant above in a real Haystack agent, with OPENAI_BASE_URL on a scripted stand-in for OpenAI's API.
  • We ran the function against a stand-in for our API that returned each answer in the table, and the MCP agent against a stand-in for our MCP server with the tools our live server lists. Through MCP, we tried each error, a 50 second answer, the long answer, a wrong tool name and the approval step.
  • We didn't try another chat generator, an OAuth sign in, a real Datacircle key or a real model.

Cost per 1,000 profiles

We charge your balance the prices on our pricing page, with no markup:

Our price per 1,000 profiles and daily limit, by provider
Up2DataHarvestAPIFetchin
Per 1,000 found$2.375$3.70$1.485
Per 1,000 not foundfree$2.30$1.485
Daily limit421 profiles per accountnonenone

Say your agent looks up 1,000 profiles in a day, one call each, and the providers find every one. You pay us at most $1.86: $1.00 for 421 through Up2Data and $0.86 for the other 579 through Fetchin. Through Fetchin alone, the same 1,000 would cost less: $1.485. Both files call Up2Data first because it bills nothing for a profile it can't find, and all our customers share Fetchin's rate limit. Your agent may call the tool more than once per question. Your model's provider bills you for its tokens.

The same tool in other agent frameworks: LinkedIn profiles in LangChain, LinkedIn profiles in Pydantic AI, LinkedIn profiles in Google ADK, LinkedIn profiles in the Claude Agent SDK, LinkedIn profiles in smolagents, LinkedIn profiles in Agno, LinkedIn profiles in Microsoft Agent Framework, LinkedIn profiles in Strands Agents and LinkedIn profiles in DSPy. The same call from a Python script: get LinkedIn profile data with Python. In Claude, ChatGPT or Cursor: our LinkedIn MCP server. Other vendors' prices per 1,000: LinkedIn profile API pricing compared.

Questions

Does Haystack have a LinkedIn tool?

Not one of its own. Bright Data's Haystack integration reads a LinkedIn profile with your Bright Data key. To read profiles through us, connect our MCP server at https://api.datacircle.dev/mcp with MCPToolset. Or write one Python function with @tool that sends POST {"url": "<the profile's LinkedIn URL>"} to https://api.datacircle.dev/v1/profiles/enrich, with the headers Authorization: Token <your key> and X-Data-Provider: up2data.

How do I send an API key to a remote MCP server in Haystack?

Pass it as the token of StreamableHttpServerInfo: StreamableHttpServerInfo(url="https://api.datacircle.dev/mcp", token=Secret.from_env_var("DATACIRCLE_API_KEY")). mcp-haystack sends it as "Authorization: Bearer <your key>" with each request to that server.

Why does my Haystack MCP tool fail after 30 seconds?

mcp-haystack waits 30 seconds by default, twice: the HTTP client's timeout in StreamableHttpServerInfo(timeout=...) and the tool call's invocation_timeout in MCPToolset or MCPTool. Raise both. Our API waits up to 45 seconds for a provider, so we set both to 60.

How much does a LinkedIn profile cost?

$2.375 per 1,000 through Up2Data (a profile it can't find is free), $3.70 per 1,000 through HarvestAPI. $1.485 per 1,000 through Fetchin, a profile it can't find billed the same. HarvestAPI bills a profile it can't find at $2.30 per 1,000.

Is there a daily limit?

Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.

Is each request live, or cached?

Live. Each request goes to the provider and gets the profile as it is today.

Do I need a LinkedIn account?

No. You send the profile's URL with your Datacircle key: no LinkedIn login, no cookies, no browser.

What happens when my balance runs out?

A call your balance can't cover answers 402. Add funds, from $5, on your dashboard.

Get started

Sign up at datacircle.dev with your work email: a $5 credit, that's 2,105 LinkedIn profiles at $2.375 per 1,000.

Sign up
Ask AI about Datacircle

Each opens with our question