Datacircle

LinkedIn profiles in DSPy: a LinkedIn tool through MCP or a Python function

You're building a program in Python with DSPy, and the model should read a LinkedIn profile from its URL: the job title, company, location and headline. dspy.ai calls DSPy "a Python framework for building AI systems", and its README says "DSPy stands for Declarative Self-improving Python." DSPy has no LinkedIn tool of its own.

Datacircle is a data co-op. Step 1: Query your favorite B2B data APIs through us. Same request, same price, no markup. Step 2: You're DONE. Every morning, you get the flat file of your data plus everyone else's. Add $50 to your account: you get $50 of API PLUS the flat file. Right now we have 3 live LinkedIn profile APIs that we trust: Up2Data, HarvestAPI and Fetchin.

You can connect our MCP server to your program and write no tool code. Or you can write one Python function that calls our API, and decide what the model reads. Both call Up2Data first. Up2Data costs $2.375 per 1,000 profiles it finds, and nothing for a profile it can't find. Fetchin is cheaper at $1.485 per 1,000, but it charges for a profile it can't find, and all our customers share its rate limit. Both switch to Fetchin only when Up2Data hits its daily limit and returns a 429.

We ran both programs in DSPy, with a scripted stand-in for OpenAI's API as the model, and a stand-in server that answers like our API. We called api.datacircle.dev with a wrong key, from the function and from the MCP code: each got a 401, at no charge. We haven't run either with a real key or a real model.

DSPy has no LinkedIn tool

A DSPy tool is a Python function: dspy.Tool reads its name, type hints and docstring, as DSPy's Tools and MCP page says. DSPy ships a few tools, such as PythonInterpreter, and converts an MCP server's tools with dspy.Tool.from_mcp_tool and LangChain's with dspy.Tool.from_langchain. None of them is a LinkedIn tool.

On October 11, 2026, GitHub's code search for "linkedin" in stanfordnlp, DSPy's GitHub organization, found 12 files. Two are in the DSPy repo: its README mentions DSPy's own page on LinkedIn, and its community resources page lists three articles people wrote on LinkedIn. The rest are in other stanfordnlp repos, such as stanza. DSPy 3.4.0, as pip installs it, holds no "linkedin" at all.

Before you start

  • Python 3.10 or later. We tested on 3.12.
  • DSPy with its MCP extra, and httpx for the function tool below:
    pip install "dspy[mcp]" httpx
    In our test, pip installed 80 packages, with mcp 2.3.0 and httpx2, the HTTP client mcp 2 uses. A plain pip install dspy left out mcp.
  • A Datacircle API key: Sign up at datacircle.dev with your work email: a $5 credit, that's 2,105 LinkedIn profiles at $2.375 per 1,000. Log in at datacircle.dev/login with your work email. Your API key is on the page once you're in. Put it in DATACIRCLE_API_KEY.
  • An OpenAI API key in OPENAI_API_KEY, for dspy.LM. DSPy's setup page says its examples "will work with any LiteLLM model string, without code changes": with another model, only the dspy.LM line changes. We haven't tried other models.

The MCP way: our MCP server through dspy.Tool.from_mcp_tool

The mcp package connects to our MCP server, and dspy.Tool.from_mcp_tool turns an MCP tool into a DSPy tool. Our MCP server is at https://api.datacircle.dev/mcp. It gets a LinkedIn profile from its URL, through Up2Data, HarvestAPI or Fetchin. This file sends your key as a Bearer token. Save it as mcp_agent.py and run python mcp_agent.py:

import asyncio
import os

import dspy
import httpx2
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client


class ReadProfile(dspy.Signature):
    """Read the LinkedIn profile at this URL with get_linkedin_profile. If up2data says its limit is reached, call it again with provider fetchin."""

    url: str = dspy.InputField()
    job_title: str = dspy.OutputField()
    company: str = dspy.OutputField()
    location: str = dspy.OutputField()
    headline: str = dspy.OutputField()


dspy.configure(lm=dspy.LM("openai/gpt-5.4-mini"))


async def main():
    key = os.environ["DATACIRCLE_API_KEY"]
    # timeout: above the 45 seconds our API waits for a provider
    async with httpx2.AsyncClient(headers={"Authorization": f"Bearer {key}"}, timeout=60) as http:
        async with streamable_http_client("https://api.datacircle.dev/mcp", http_client=http) as (read, write):
            async with ClientSession(read, write) as session:
                await session.initialize()
                listed = await session.list_tools()
                tools = [dspy.Tool.from_mcp_tool(session, tool) for tool in listed.tools if tool.name == "get_linkedin_profile"]
                agent = dspy.ReAct(ReadProfile, tools=tools, max_iters=4)  # at most 4 tool calls, so at most 4 billed requests to our API
                result = await agent.acall(url="https://www.linkedin.com/in/example-profile")
                print(result.job_title, result.company, result.location, result.headline, sep="\n")


asyncio.run(main())

The three async with blocks open the HTTP client, the connection and the session, and close them at the end. Call the agent with await agent.acall(...) inside them. With agent(...), no request reached our server, and the model read a ValueError telling it to use acall.

DSPy's MCP tutorial shows a local server only. An older DSPy page connects over HTTP with streamablehttp_client, and on mcp 2.3.0 that import failed:

ImportError: cannot import name 'streamablehttp_client' from 'mcp.client.streamable_http' (…). Did you mean: 'streamable_http_client'?

Its other example, Client(url), takes no headers: our server answered with a 401 before any model call. With this file's HTTP client, Client(streamable_http_client(url, http_client=http)) worked against our stand-in server.

The file keeps only get_linkedin_profile of our server's tools. Check its spelling: with a name our server doesn't have, the agent got no tool and no error. The first request to the model was 4,107 bytes with one tool, and 6,119 with all of them.

DSPy gives the model our JSON text, uncut

Our server sends each answer as text and as structuredContent. DSPy gives the model the text alone: the provider's JSON, plus datacircle_meta, what the call cost and your balance after it:

{"data": {…}, "meta": {…}, "datacircle_meta": {…}}

Up2Data's example answer was 1,167 characters. DSPy passed a saved Fetchin answer of 65,229 characters to the model whole. DSPy leaves out the tool's output schema, all 8,291 characters of it.

At Up2Data's limit, the model reads our error at the end of a traceback

The tool takes url, and provider: up2data (the default), harvestapi or fetchin. At Up2Data's limit, the server tells your agent to call again through Fetchin or HarvestAPI. HarvestAPI costs more per 1,000 profiles than the other two, so we name Fetchin in the file's signature. DSPy raises on our error, and dspy.ReAct gives the model a Python traceback with our JSON at its end:

Execution error in get_linkedin_profile:
Traceback (most recent call last):
  …
RuntimeError: Failed to call a MCP tool: {"error": "daily up2data limit reached for your account …"}

The traceback holds no status code. Our scripted model read it, called again with fetchin and got the profile. Every other error reached the model as a traceback, and the agent kept going. Our MCP server docs list every tool.

With a wrong key, your program stops at session.initialize()

With a wrong key, session.initialize() raised before any model call. We tried it on api.datacircle.dev:

  + Exception Group Traceback (most recent call last):
  …
      |     await session.initialize()
  …
      | mcp.shared.exceptions.MCPError: invalid API key or access token

Our error message is the last line. Set the right key in DATACIRCLE_API_KEY.

Set the HTTP client's timeout above the 45 seconds our API waits for a provider

DSPy sets no time limit on an MCP call: you set it on the HTTP client you pass, and an httpx2.AsyncClient waits 5 seconds by default. Our stand-in server waited 50 seconds before it answered. With timeout=60, the model got the profile. Without a timeout of your own, the program stopped after 5 seconds on an httpx2.ReadTimeout, and the model never read the profile.

To cut a slow call and go on, pass read_timeout_seconds to ClientSession. We set it to 10, and after 10 seconds the model read mcp.shared.exceptions.MCPError: Request 'tools/call' timed out, and the agent kept going.

The function way: one Python function that calls our API

You send the provider's own request to api.datacircle.dev, with your Datacircle key. That's the only change. The function sends Up2Data's and Fetchin's own requests, with X-Data-Provider naming the provider. Up2Data's request is in our API reference. Save this file as tool_agent.py and run python tool_agent.py:

import os
import time

import dspy
import httpx

API = os.environ.get("DATACIRCLE_API_URL", "https://api.datacircle.dev")
KEY = os.environ["DATACIRCLE_API_KEY"]


def call(provider, method, path, **request):
    """One call to Datacircle's API through one provider. Fetchin's 429 is its rate limit: wait a second and send it again."""
    for wait in (0, 1, 2):
        time.sleep(wait)
        answer = httpx.request(method, f"{API}{path}", headers={"Authorization": f"Token {KEY}", "X-Data-Provider": provider}, timeout=60, **request)
        if provider != "fetchin" or answer.status_code != 429:
            return answer
    return answer


def get_linkedin_profile(url: str) -> dict:
    """Get the current job title, company, location and headline on a LinkedIn profile, from the profile's URL,
    like https://www.linkedin.com/in/example-profile.
    Returns the keys job_title, company, location and headline, or one key, error, when the profile can't be read.
    Each call asks the provider live and may be billed to the Datacircle balance."""
    answer = call("up2data", "POST", "/v1/profiles/enrich", json={"url": url})
    if answer.status_code == 200:
        profile = answer.json()["data"]
        company = profile.get("current_company") or {}
        return {"job_title": company.get("title"), "company": company.get("name"),
                "location": (profile.get("location") or {}).get("raw"), "headline": profile.get("headline")}
    if answer.status_code == 429:  # Up2Data's daily limit: Fetchin answers instead
        answer = call("fetchin", "GET", "/api/v1/profile", params={"profileUrlOrUrn": url})
        if answer.status_code == 200:
            profile = answer.json()
            return {"job_title": profile.get("jobTitle"), "company": profile.get("companyName"),
                    "location": profile.get("location"), "headline": profile.get("title")}
    if answer.status_code in (404, 422):
        return {"error": "This LinkedIn profile is private or deleted."}
    if answer.status_code == 400:
        return {"error": "This is not a LinkedIn profile URL. Send one like https://www.linkedin.com/in/example-profile"}
    if answer.status_code == 402:
        return {"error": "The Datacircle balance is too low for this call. Tell the user to add funds on their Datacircle dashboard."}
    if answer.status_code in (429, 500, 502, 503, 504):
        return {"error": f"The provider didn't answer ({answer.status_code}), and the call wasn't charged. Try again in a minute."}
    answer.raise_for_status()  # 401: DATACIRCLE_API_KEY is wrong


class ReadProfile(dspy.Signature):
    """Read the LinkedIn profile at this URL."""

    url: str = dspy.InputField()
    job_title: str = dspy.OutputField()
    company: str = dspy.OutputField()
    location: str = dspy.OutputField()
    headline: str = dspy.OutputField()


dspy.configure(lm=dspy.LM("openai/gpt-5.4-mini"))
agent = dspy.ReAct(ReadProfile, tools=[get_linkedin_profile], max_iters=4)  # at most 4 tool calls, so at most 4 billed requests to our API
result = agent(url="https://www.linkedin.com/in/example-profile")
print(result.job_title, result.company, result.location, result.headline, sep="\n")

dspy.ReAct turns the function into a dspy.Tool, as DSPy's ReAct page shows. The model gets the function's name, its whole docstring, and url as a string, but not -> dict, so the docstring names the answer's keys.

The function sends the URL to Up2Data. At Up2Data's daily limit (its 429), it sends the same URL to Fetchin. Fetchin takes 5 requests a second across all our customers, and returns a 429 past that: the function waits one second and sends it again, then waits two seconds and sends it once more.

The function returns a dict, and DSPy gives it to the model as JSON text:

{"job_title": …, "company": …, "location": …, "headline": …}

For a private or deleted profile, a URL that isn't a profile, a low balance or a provider error, the function returns one key, error, so the agent can tell the user and go on. Each field comes from the same JSON path as in our Python post:

Source of each field
FieldUp2Data's answerFetchin's answer
job_titledata.current_company.titlejobTitle
companydata.current_company.namecompanyName
locationdata.location.rawlocation
headlinedata.headlinetitle

With a wrong key (401), the function raises. DSPy prints nothing, gives the model the error, and the agent keeps going. We tried a wrong key on api.datacircle.dev, and the model read this:

Execution error in get_linkedin_profile:
Traceback (most recent call last):
  …
httpx.HTTPStatusError: Client error '401 Unauthorized' for url 'https://api.datacircle.dev/v1/profiles/enrich'

Set the right key in DATACIRCLE_API_KEY. DSPy sets no time limit on a function, and httpx waits 60 seconds here: our stand-in server waited 50 seconds, and the model got the answer. DATACIRCLE_API_URL is for tests: point it at a stand-in for our API, and you can run the function without spending your balance.

dspy.ReAct writes the tools into the prompt, even with dspy.ChatAdapter(use_native_function_calling=True). DSPy's ReAct and ReActV2 page says the experimental dspy.ReActV2 becomes dspy.ReAct in DSPy 3.5. With dspy.ReActV2 and that adapter, DSPy sent the function to the model as an OpenAI tool, and its result as a tool message.

Cap the agent's calls with max_iters

Each tool call reaches our API, and we may bill it. We bill each call that gets a profile, and each call for a profile Fetchin or HarvestAPI can't find, at the price on our pricing page. The function tool sends up to four requests per call, and we bill at most one of them: Up2Data's 429 and Fetchin's 429 are free.

dspy.ReAct stops after max_iters steps, 20 by default: a step is one call to the model and the tool it picks. A scripted model that asked for the tool at every step made 20 calls to a stand-in for our API in under a second. With max_iters=4, as in both files, it got 4 profiles. At the limit, DSPy raises no error and logs nothing, and asks the model for the four fields. In result.trajectory, the last tool name is finish only when the model stopped on its own.

At each step, DSPy sends the model the whole trajectory so far. With the long Fetchin answer at each step, the requests grew from 4,107 bytes to 73,207, 142,311 and 211,415, then 277,648 for the fields. DSPy's ReAct page says it drops the oldest step when a request is too long for the model.

DSPy caches the model's answers on disk. On a second run with the same URL, the function file made no model call and still called our API: we bill both runs.

DSPy has no step that asks you before a tool call. If a callback raised an error before the call, DSPy logged a warning, and our server still got the call.

API answers: cost and what the tool returns

Our API's answers to the function in tool_agent.py
AnswerMeaningCostThe tool returns
Up2Data 200the profile$2.375 per 1,000the four fields
Up2Data 422the profile is private or deletedfreeerror: "This LinkedIn profile is private or deleted."
Up2Data 400not a LinkedIn profile URLfreeerror: "This is not a LinkedIn profile URL."
Up2Data 429its daily limitfreeFetchin's answer
Fetchin 200the profile$1.485 per 1,000the four fields
Fetchin 404 with PROFILE_NOT_FOUNDthe profile is private or deleted$1.485 per 1,000: Fetchin bills the lookuperror: "This LinkedIn profile is private or deleted."
Fetchin 429its rate limit, which all our customers sharefreeFetchin's answer after up to two retries, or error: "Try again in a minute"
402your balance can't cover the callfreeerror: "The Datacircle balance is too low for this call."
500, 502, 503 or 504the provider failed, or didn't answer within 45 secondsfreeerror: "Try again in a minute"
401your key is wrongfreean exception, which the model reads as a traceback after Execution error in get_linkedin_profile

Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.

The tests we ran

  • We tested on October 11, 2026, with Python 3.12, dspy 3.4.0, litellm 1.105.0, mcp 2.3.0, httpx2 2.13.1, openai 2.54.0 and httpx 0.28.1.
  • We ran each file and each variant above in DSPy, with OPENAI_BASE_URL on a scripted stand-in for OpenAI's API. With it set, DSPy sends the model's requests through LiteLLM, and without it, through its own client.
  • We ran the function against a stand-in for our API that returned each answer in the table. We ran the MCP agent against a stand-in for our MCP server with the tools our live server lists, and tried each error, a 50 second answer, both timeouts, the long answer and a wrong tool name.
  • We didn't try an OAuth sign in, a real Datacircle key or a real model.

Cost per 1,000 profiles

We charge your balance the prices on our pricing page, with no markup:

Our price per 1,000 profiles and daily limit, by provider
Up2DataHarvestAPIFetchin
Per 1,000 found$2.375$3.70$1.485
Per 1,000 not foundfree$2.30$1.485
Daily limit421 profiles per accountnonenone

Say your agent looks up 1,000 profiles in a day, one call each, and the providers find every one. You pay us at most $1.86: $1.00 for 421 through Up2Data and $0.86 for the other 579 through Fetchin. Through Fetchin alone, the same 1,000 would cost less: $1.485. Both files call Up2Data first because it bills nothing for a profile it can't find, and all our customers share Fetchin's rate limit. Your agent may call the tool more than once per question. Your model's provider bills you for its tokens.

The same tool in other agent frameworks: LinkedIn profiles in LangChain, LinkedIn profiles in Pydantic AI, LinkedIn profiles in Google ADK, LinkedIn profiles in the Claude Agent SDK, LinkedIn profiles in smolagents, LinkedIn profiles in Agno, LinkedIn profiles in Microsoft Agent Framework, LinkedIn profiles in Strands Agents and LinkedIn profiles in Haystack. The same call from a script: get LinkedIn profile data with Python. In Claude, ChatGPT or Cursor: our LinkedIn MCP server. Other vendors' prices per 1,000: LinkedIn profile API pricing compared.

Questions

Does DSPy have a LinkedIn tool?

Not one of its own. Connect our MCP server at https://api.datacircle.dev/mcp and convert its tool with dspy.Tool.from_mcp_tool, or give dspy.ReAct one Python function that sends POST {"url": "<the profile's LinkedIn URL>"} to https://api.datacircle.dev/v1/profiles/enrich, with the headers Authorization: Token <your key> and X-Data-Provider: up2data.

How do I send an API key to a remote MCP server in DSPy?

DSPy leaves the connection to the mcp package. With mcp 2, make an httpx2.AsyncClient with the header Authorization: Bearer <your key> and a timeout. Pass it to streamable_http_client as http_client, open a ClientSession on it, and convert each tool with dspy.Tool.from_mcp_tool(session, tool).

Why can't I import streamablehttp_client for DSPy?

pip installs mcp 2 today, which renamed it streamable_http_client. On mcp 2.3.0, importing the old name raised ImportError in our test.

What does max_iters default to in dspy.ReAct?

20 in DSPy 3.4.0: up to 20 tool calls in one run. At the limit, DSPy asks the model for the outputs from the tool results so far, with no error or warning.

How much does a LinkedIn profile cost?

$2.375 per 1,000 through Up2Data (a profile it can't find is free), $3.70 per 1,000 through HarvestAPI. $1.485 per 1,000 through Fetchin, a profile it can't find billed the same. HarvestAPI bills a profile it can't find at $2.30 per 1,000.

Is there a daily limit?

Up2Data takes $1 a day per account (421 profiles), with a shared daily limit for all customers, then answers 429 until 00:00 UTC. HarvestAPI has no daily limit. Fetchin has no daily limit either.

Is each request live, or cached?

Live. Each request goes to the provider and gets the profile as it is today.

Do I need a LinkedIn account?

No. You send the profile's URL with your Datacircle key: no LinkedIn login, no cookies, no browser.

What happens when my balance runs out?

A call your balance can't cover answers 402. Add funds, from $5, on your dashboard.

Get started

Sign up at datacircle.dev with your work email: a $5 credit, that's 2,105 LinkedIn profiles at $2.375 per 1,000.

Sign up
Ask AI about Datacircle

Each opens with our question