Automationscribe.com
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us
No Result
View All Result
Automation Scribe
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us
No Result
View All Result
Automationscribe.com
No Result
View All Result

Device Calling vs. Code Execution for AI Brokers: Selecting the Proper Motion Primitive

admin by admin
October 5, 2026
in Artificial Intelligence
0
Device Calling vs. Code Execution for AI Brokers: Selecting the Proper Motion Primitive
399
SHARES
2.3k
VIEWS
Share on FacebookShare on Twitter


On this article, you’ll be taught what device calling and code execution are as agent motion primitives, how they differ mechanically, and when to decide on one over the opposite.

Matters we’ll cowl embrace:

  • How device calling works below the hood, and why it stays the fitting alternative for single, time-sensitive lookups.
  • How code execution by way of Programmatic Device Calling differs from normal device calling, and what measurable advantages it provides for fan-out and aggregation duties.
  • A sensible choice framework for selecting between the 2 primitives based mostly on name depend, information sensitivity, latency, infrastructure, and auditability wants.

Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive

Image an agent asking one simple-sounding query: which of twenty staff went over their Q3 journey funds. To reply it, the agent wants every particular person’s expense line objects, each flight, resort, and meal receipt, in contrast in opposition to a funds restrict tied to their stage. Constructed the plain means, with the mannequin calling a device for every particular person’s bills one by one, that’s twenty separate device calls, every returning fifty to 100 line objects, and each single a kind of objects has to go via the mannequin’s context simply so it may be added up. That’s over 2,000 line objects and greater than 50KB of uncooked information the mannequin by no means really wanted to learn — it wanted a sum.

That’s the true price hiding behind a design choice most agent tutorials skip previous completely: how does an agent really take motion on the earth. There are two actual solutions, device calling and code execution, and which one you attain for isn’t a method choice — it’s an architectural alternative with measurable penalties for price, latency, and accuracy. This text breaks down each motion primitives for AI brokers intimately, builds an actual, runnable instance of every utilizing the identical underlying device, and closes with an sincere, numbers-backed framework for selecting between them. In case you haven’t constructed a fundamental tool-calling agent but, take a look at this text, Simple Agentic Device Calling with Gemma 4 — it’s the pure place to begin earlier than this one.

What Is an Motion Primitive, and Why Does the Selection Matter?

An motion primitive is the basic mechanism by which a language mannequin turns a choice into an actual impact on the earth — a database write, an API name, a file learn. Each agent framework, no matter else it does, is constructed on high of certainly one of these primitives at its core.

Device calling is the primitive most individuals be taught first: the mannequin produces one structured request at a time, a number utility executes it, and the consequence comes again into the dialog earlier than the mannequin decides what to do subsequent. Code execution is the newer different: as an alternative of requesting one motion and ready, the mannequin writes an precise program — in Python or TypeScript — that performs a number of actions in sequence or in parallel, and solely this system’s ultimate output returns to the mannequin.

Neither one is a wrapper across the different, and neither has quietly changed the opposite. They’re genuinely totally different mechanisms with totally different failure modes, totally different infrastructure necessities, and totally different price profiles, and the remainder of this text is about understanding each properly sufficient to choose appropriately.

Device Calling

It’s price understanding what’s really occurring beneath a device name, as a result of the mechanics clarify each its strengths and its actual limitations. Based on Cloudflare’s detailed breakdown of the method, a mannequin producing a device name doesn’t produce strange textual content. It’s been particularly educated to output a pair of particular tokens — one signaling “the next is a device name” and one other marking its finish — with a JSON payload describing the device identify and arguments sitting between them. The appliance operating the mannequin watches for these tokens, pauses era the second it sees the closing one, parses the JSON in opposition to a schema you outlined, really executes the decision, and feeds the consequence again into the dialog as if it had been the subsequent factor the consumer mentioned.

That’s a clear, auditable, one-step-at-a-time loop, and it’s precisely why device calling grew to become the default. Each motion is a discrete, loggable occasion. Each result’s one thing the mannequin straight sees and may cause about in pure language earlier than deciding what occurs subsequent.

Code Execution

Code execution takes a unique beginning place completely: as an alternative of asking the mannequin to explain an motion in a constrained JSON format, you let it write precise code that performs the motion, operating in a sandboxed setting separate from the mannequin itself. Anthropic’s unique code-execution-with-MCP sample frames this exactly as presenting your instruments as a code API quite than a set of straight callable features, so the mannequin can write a script that imports precisely the instruments it wants and calls them the way in which it might name every other perform.

The mechanism that makes this genuinely totally different — not only a relabeled device name — is what Anthropic now calls Programmatic Device Calling, launched alongside two companion options in November 2025. Reasonably than every device consequence flowing again via the mannequin one by one, you mark particular instruments as callable from code by including an allowed_callers discipline to their definition, and add a code_execution device to the request. When the mannequin needs to behave, it writes a full script — loops, conditionals, error dealing with, and all — that calls these instruments straight inside a sandboxed execution setting. Every particular person device name the script makes nonetheless executes precisely the way in which it might in strange device calling; you continue to obtain a request and return a consequence, however that result’s intercepted and processed by the operating script quite than being pushed into the mannequin’s context. Solely when the script finishes does its ultimate output — and nothing else — return to the mannequin.

That’s your entire distinction in a single sentence: device calling places each intermediate end in entrance of the mannequin; code execution lets the mannequin resolve, via the code it writes, precisely what makes it again.

A side-by-side flow diagram of Tool Calling and Code Execution

A side-by-side move diagram of Device Calling and Code Execution (click on to enlarge)

Device Calling for a Single, Time-Delicate Lookup

Principle is simpler to belief as soon as it’s operating in opposition to an actual API, so each examples on this article use the identical device — a get_weather perform backed by Open-Meteo, a free climate API that wants no API key in any respect, solely an Anthropic API key to run the agent itself.

Begin with the case device calling is clearly proper for: a single query that wants one lookup and a natural-language reply — “what’s the climate like in London proper now.”

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

25

26

27

28

29

30

31

32

33

34

35

36

37

38

39

40

41

42

43

44

45

46

47

48

49

50

51

52

53

54

55

56

57

58

59

60

61

62

63

64

65

66

67

68

69

70

71

72

73

74

75

76

77

78

79

80

81

82

83

84

85

86

import json

import requests

from anthropic import Anthropic

 

shopper = Anthropic()  # reads ANTHROPIC_API_KEY from the setting

 

def get_weather(metropolis: str) -> dict:

    “”“Search for a metropolis’s coordinates, then fetch its present temperature

    and this week’s each day highs from Open-Meteo’s free, keyless API.”“”

    geo = requests.get(

        “https://geocoding-api.open-meteo.com/v1/search”,

        params={“identify”: metropolis, “depend”: 1},

    ).json()

    if not geo.get(“outcomes”):

        return {“error”: f“Couldn’t discover a location named ‘{metropolis}'”}

    lat = geo[“results”][0][“latitude”]

    lon = geo[“results”][0][“longitude”]

 

    forecast = requests.get(

        “https://api.open-meteo.com/v1/forecast”,

        params={

            “latitude”: lat,

            “longitude”: lon,

            “present”: “temperature_2m”,

            “each day”: “temperature_2m_max”,

            “timezone”: “auto”,

        },

    ).json()

 

    return {

        “metropolis”: metropolis,

        “current_temp_c”: forecast[“current”][“temperature_2m”],

        “week_high_temps_c”: forecast[“daily”][“temperature_2m_max”],

        “unit”: “celsius”,

    }

 

weather_tool = {

    “identify”: “get_weather”,

    “description”: (

        “Get the present temperature and this week’s each day excessive “

        “temperatures for a metropolis. Returns JSON with metropolis, “

        “current_temp_c, week_high_temps_c (7 each day highs), and unit.”

    ),

    “input_schema”: {

        “kind”: “object”,

        “properties”: {

            “metropolis”: {“kind”: “string”, “description”: “Metropolis identify, e.g. ‘Lagos'”}

        },

        “required”: [“city”],

    },

}

 

messages = [{“role”: “user”, “content”: “What’s the weather like in London right now?”}]

 

response = shopper.messages.create(

    mannequin=“claude-sonnet-5”,

    max_tokens=1024,

    instruments=[weather_tool],

    messages=messages,

)

 

# Preserve resolving device calls till Claude produces a ultimate textual content reply

whereas response.stop_reason == “tool_use”:

    messages.append({“position”: “assistant”, “content material”: response.content material})

    tool_results = []

 

    for block in response.content material:

        if block.kind == “tool_use” and block.identify == “get_weather”:

            consequence = get_weather(**block.enter)

            tool_results.append({

                “kind”: “tool_result”,

                “tool_use_id”: block.id,

                “content material”: json.dumps(consequence),

            })

 

    messages.append({“position”: “consumer”, “content material”: tool_results})

    response = shopper.messages.create(

        mannequin=“claude-sonnet-5”,

        max_tokens=1024,

        instruments=[weather_tool],

        messages=messages,

    )

 

for block in response.content material:

    if block.kind == “textual content”:

        print(block.textual content)

Strolling via what issues right here: get_weather itself is strange Python — nothing agent-specific about it — it geocodes a metropolis identify and pulls each the present temperature and the week’s each day highs in a single request. The weather_tool dictionary is the schema Claude really sees, and the outline issues greater than it seems — a obscure description is without doubt one of the most typical causes of a mannequin calling a device with the flawed arguments. The whereas response.stop_reason == “tool_use” loop is the true mechanical coronary heart of ordinary device calling: each time Claude requests the device, your code has to truly run it, wrap the consequence as a tool_result block, append it to the dialog, and name the API once more — and this repeats for as many device calls as the duty wants. For a single lookup like this one, that’s one go via the loop and accomplished, which is strictly why device calling matches this case properly: one name, one consequence, and a consequence small and related sufficient that Claude genuinely advantages from seeing it straight earlier than writing a natural-language reply.

Code Execution for Fan-Out and Aggregation

Now change the query, utilizing the very same get_weather perform — utterly unchanged: “given these fifteen cities, which one can have the coldest excessive temperature this week, and what’s the common weekly excessive throughout all of them?”

Run that via the tool-calling loop above and also you’d get fifteen separate device calls, fifteen full JSON payloads of each day temperatures pushed into Claude’s context, and Claude would then need to manually evaluate and common them in pure language — gradual, token-expensive, and precisely the sort of arithmetic a mannequin is extra error-prone at than a for-loop is. That is exactly the case Programmatic Device Calling was constructed for.

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

25

26

27

28

29

30

31

32

33

34

35

36

37

38

39

40

41

42

43

44

45

46

47

48

49

50

51

52

53

54

55

56

57

58

59

60

61

62

63

64

65

66

67

68

69

70

71

72

73

74

75

76

77

78

79

80

81

82

import json

from anthropic import Anthropic

 

shopper = Anthropic()

 

# Similar get_weather perform from the earlier instance, unchanged

 

weather_tool = {

    “identify”: “get_weather”,

    “description”: (

        “Get the present temperature and this week’s each day excessive “

        “temperatures for a metropolis. Returns JSON with metropolis, “

        “current_temp_c, week_high_temps_c (7 each day highs), and unit.”

    ),

    “input_schema”: {

        “kind”: “object”,

        “properties”: {

            “metropolis”: {“kind”: “string”, “description”: “Metropolis identify, e.g. ‘Lagos'”}

        },

        “required”: [“city”],

    },

    # That is the one line that adjustments the primitive: it opts the device

    # into being known as from inside generated code, not solely straight

    # by the mannequin

    “allowed_callers”: [“code_execution_20250825”],

}

 

code_execution_tool = {“kind”: “code_execution_20250825”, “identify”: “code_execution”}

 

cities = [

    “Lagos”, “Nairobi”, “Cairo”, “Accra”, “Kigali”,

    “Casablanca”, “Addis Ababa”, “Dakar”, “Tunis”, “Kampala”,

    “Harare”, “Lusaka”, “Maputo”, “Windhoek”, “Gaborone”,

]

 

messages = [{

    “role”: “user”,

    “content”: (

        f“Given these cities: {‘, ‘.join(cities)}, which one will have “

        “the coldest high temperature this week, and what’s the average “

        “weekly high across all of them? Use the get_weather tool.”

    ),

}]

 

response = shopper.beta.messages.create(

    betas=[“advanced-tool-use-2025-11-20”],

    mannequin=“claude-sonnet-5”,

    max_tokens=2048,

    instruments=[code_execution_tool, weather_tool],

    messages=messages,

)

 

# The loop seems just like normal device calling, however now some

# tool_use blocks carry a “caller” discipline, that means the request got here

# from inside Claude’s generated script quite than from Claude straight

whereas response.stop_reason == “tool_use”:

    messages.append({“position”: “assistant”, “content material”: response.content material})

    tool_results = []

 

    for block in response.content material:

        if block.kind == “tool_use” and block.identify == “get_weather”:

            consequence = get_weather(**block.enter)

            tool_results.append({

                “kind”: “tool_result”,

                “tool_use_id”: block.id,

                “content material”: json.dumps(consequence),

            })

 

    if tool_results:

        messages.append({“position”: “consumer”, “content material”: tool_results})

 

    response = shopper.beta.messages.create(

        betas=[“advanced-tool-use-2025-11-20”],

        mannequin=“claude-sonnet-5”,

        max_tokens=2048,

        instruments=[code_execution_tool, weather_tool],

        messages=messages,

    )

 

for block in response.content material:

    if block.kind == “textual content”:

        print(block.textual content)

The one most vital line on this complete script is “allowed_callers”: [“code_execution_20250825”]. With out it, the device behaves precisely because it did within the earlier instance — callable solely straight by the mannequin. With it added, Claude positive factors the choice to jot down a script that calls get_weather fifteen instances itself, doubtless in parallel utilizing asyncio.collect, sum and kind the outcomes, and print solely the ultimate reply — the coldest metropolis and the common — to straightforward output. Your Python code doesn’t change the way it responds to particular person device calls in any respect; that a part of the loop seems almost similar to the tool-calling instance. What adjustments is invisible out of your facet of the API: fourteen of these fifteen climate lookups, and each intermediate comparability between them, by no means contact Claude’s context.

Claude solely ever sees the 2 numbers it really requested for. Since this makes use of a beta function, it’s price double-checking the precise beta header string and block-handling particulars in opposition to Anthropic’s present documentation earlier than counting on it in manufacturing, as beta APIs are the a part of any platform almost certainly to shift.

Why Code Execution Wins at Scale

The climate instance makes the mechanism seen, nevertheless it’s price backing this up with actual, printed figures quite than instinct alone. Anthropic’s unique code-execution-with-MCP sample took an actual Google Drive-to-Salesforce workflow from 150,000 tokens right down to 2,000 — a 98.7% discount — just by holding a full assembly transcript contained in the execution setting as an alternative of routing it via the mannequin twice.

Programmatic Device Calling’s personal inner benchmarking, reported straight by Anthropic, discovered common token utilization on complicated analysis duties dropped from 43,588 to 27,297 — a 37% discount — whereas accuracy on the GAIA benchmark really improved, rising from 46.5% to 51.2%, and inner data retrieval accuracy rose from 25.6% to twenty-eight.5%. That final element issues greater than the token financial savings alone: this isn’t purely a price optimization. Offloading orchestration logic to precise code quite than asking a mannequin to trace it via pure language measurably reduces the sort of errors that come from a mannequin shedding monitor of a dozen intermediate values it’s making an attempt to check in its head.

The tutorial consequence beneath all of this predates Anthropic’s personal tooling. The unique CodeAct paper from Wang and colleagues in 2024 discovered that brokers taking motion via executable code, quite than JSON-formatted device calls, succeeded as much as 20% extra usually on complicated, multi-step duties. Code execution isn’t a current product function bolted onto an current thought — it’s a research-backed sample that the main labs have spent the previous two years turning into manufacturing infrastructure.

The place Device Calling Nonetheless Wins

The numbers above could make code execution seem like an unconditional improve, and it isn’t one. There’s an actual, sincere case for sticking with plain device calling in a significant set of conditions.

Single-call duties are the clearest case. The Lagos climate instance earlier on this article positive factors nothing from a sandbox — one name, one small consequence, and the overhead of spinning up a code execution setting provides latency with out including any actual profit. Duties the place the mannequin genuinely must cause over an intermediate end in pure language are the second case: if the precise level of a step is for the mannequin to note one thing delicate in a doc or a dataset and reply to it conversationally, filtering that information away in a sandbox defeats the aim. Easier infrastructure is an actual, sensible issue too — a workforce with out an current safe sandboxing setup takes on actual operational price standing one up, and that price must be weighed in opposition to the financial savings, not assumed away. And auditability issues greater than it will get credit score for: a device name is one clear, loggable occasion with a reputation and a set of arguments, whereas reasoning about precisely what a generated script did internally — particularly after the very fact, throughout an incident — is a genuinely more durable debugging downside.

Determination Framework: Selecting the Proper Primitive

Pulling every thing above into one sensible reference:

Issue Favors device calling Favors code execution
Variety of calls wanted One, or a small, fastened few A number of, particularly with fan-out or aggregation
What occurs to outcomes The mannequin must learn and cause over them straight They simply have to be filtered, summed, or in contrast
Information sensitivity Low — nothing problematic concerning the mannequin seeing it Excessive — PII or massive payloads higher saved out of context
Latency tolerance Tight — sandbox startup isn’t price paying for Workflow already entails a number of round-trips anyway
Group infrastructure No current sandboxing setup Sandbox or code-execution tooling already in place
Auditability wants Each discrete motion have to be individually logged Mixture final result issues greater than every inner step

The Hybrid Actuality: Most Manufacturing Brokers Use Each

It’s price closing this out by pushing again gently on the framing of the article’s personal title. In observe, this isn’t a everlasting, once-and-for-all architectural choice — it’s a per-task judgment name, and Anthropic’s personal steering treats it precisely that means. Their superior device use launch shipped Programmatic Device Calling alongside two companion options particularly meant to be layered collectively as wanted: a Device Search Device for locating the fitting device out of a giant library with out loading each definition upfront, and Device Use Examples for instructing a mannequin the conventions a schema alone can’t specific. Their very own suggestion is to begin with whichever bottleneck is definitely limiting a given agent — context bloat from too many device definitions, massive intermediate outcomes polluting context, or parameter errors — and add the matching function, quite than reaching for each functionality on day one.

A single well-built agent, in observe, tends to make use of plain device calling for its easy, single-shot lookups and swap to code execution the second a activity requires fan-out, aggregation, or dealing with information too massive or delicate to place in entrance of the mannequin straight. The precise ability price constructing isn’t selecting a primitive as soon as — it’s recognizing, activity by activity, which one the work in entrance of you really wants.

Conclusion

An motion primitive is infrastructure, not a choice, and the 2 examples constructed on this article show it with the identical fifteen strains of device definition beneath each. Get it proper and an agent handles a fan-out activity throughout fifteen cities — or two thousand expense line objects — in a single clear go. Get it flawed — attain for device calling on a activity that wants code execution — and nothing crashes. The agent nonetheless solutions. It simply does it slower, extra expensively, and with a context window quietly stuffed with information no one really wanted to learn.

Tags: ActionAgentsCallingChoosingcodeExecutionPrimitivetool
Previous Post

Govern AI Brokers

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Popular News

  • Greatest practices for Amazon SageMaker HyperPod activity governance

    Greatest practices for Amazon SageMaker HyperPod activity governance

    406 shares
    Share 162 Tweet 102
  • How Cursor Really Indexes Your Codebase

    405 shares
    Share 162 Tweet 101
  • Construct a serverless audio summarization resolution with Amazon Bedrock and Whisper

    404 shares
    Share 162 Tweet 101
  • Introducing the Amazon Bedrock AgentCore Code Interpreter

    404 shares
    Share 162 Tweet 101
  • Context Engineering — A Complete Fingers-On Tutorial with DSPy

    404 shares
    Share 162 Tweet 101

About Us

Automation Scribe is your go-to site for easy-to-understand Artificial Intelligence (AI) articles. Discover insights on AI tools, AI Scribe, and more. Stay updated with the latest advancements in AI technology. Dive into the world of automation with simplified explanations and informative content. Visit us today!

Category

  • AI Scribe
  • AI Tools
  • Artificial Intelligence

Recent Posts

  • Device Calling vs. Code Execution for AI Brokers: Selecting the Proper Motion Primitive
  • Govern AI Brokers
  • Advantageous-tune a search agent with multi-turn RL on Amazon SageMaker AI
  • Home
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms & Conditions

© 2024 automationscribe.com. All rights reserved.

No Result
View All Result
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us

© 2024 automationscribe.com. All rights reserved.