Automationscribe.com
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us
No Result
View All Result
Automation Scribe
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us
No Result
View All Result
Automationscribe.com
No Result
View All Result

Constructing a context-aware AI assistant on AgentCore and OpenClaw

admin by admin
October 7, 2026
in Artificial Intelligence
0
Constructing a context-aware AI assistant on AgentCore and OpenClaw
399
SHARES
2.3k
VIEWS
Share on FacebookShare on Twitter


Off-the-shelf AI assistants reply particular person questions nicely, however they fall brief on a unique axis: continuity. Ask a stateless assistant about your backyard at present and it has no thought that you simply talked about your fast-draining raised beds three weeks in the past, that you simply solely use natural fertilizer, or that your petunias had been struggling by a warmth wave. Each dialog begins from zero, and the burden of re-explaining context falls on the person.

The issue isn’t the standard of the solutions, however that the assistant has no reminiscence of you. This publish exhibits how you can construct a private assistant that accumulates context utilizing OpenClaw, an open supply agentic system, operating on AgentCore runtime, a functionality of Amazon Bedrock AgentCore. AgentCore reminiscence, a functionality of Amazon Bedrock AgentCore, turns disposable chats into sturdy data. Additionally, you will see how you can tag these reminiscences with structured metadata to retrieve data that matter for the query at hand.

Our operating instance is Sprout, a gardening assistant, however the structure is domain-agnostic. Swap the persona and the talents manifest, and the identical pipeline serves a assist bot, a health coach, or an inside assist desk. The complete system lives in a single AWS CloudFormation template, deploys with one command, and runs on a consumption-based mannequin that prices just a few {dollars} a month for mild private use. Alongside the best way, we share design tips you possibly can apply to assistants you construct on this stack.

Resolution overview

AgentCore is a platform to construct, join, and optimize brokers at scale, with any framework or mannequin. The next diagram exhibits the end-to-end request circulate, from an inbound Telegram webhook by the AgentCore runtime, and its supporting AWS companies.

Determine 1: Telegram webhooks and Amazon EventBridge schedules each invoke the identical AgentCore runtime agent, which coordinates the OpenClaw gateway, AgentCore reminiscence, and Amazon Bedrock

Two entry factors converge on one agent. Telegram messages arrive by Amazon API Gateway and a webhook AWS Lambda perform, whereas scheduled jobs corresponding to morning watering reminders arrive by Amazon EventBridge Scheduler and a cronjob Lambda perform. Each name the InvokeAgentRuntime API on the AgentCore runtime, the place a skinny server.py course of coordinates the OpenClaw gateway, AgentCore reminiscence, and the Amazon Bedrock Converse API. Amazon Easy Storage Service (Amazon S3) offers workspace storage, AWS Key Administration Service (AWS KMS) handles encryption, AWS Secrets and techniques Supervisor holds the bot token, and Amazon CloudWatch captures logs and metrics.

Stipulations

To deploy your personal model utilizing the Launch Stack button or scripts/deploy.sh (described within the Develop your personal part), you have to:

  • Amazon Bedrock AgentCore entry, together with AgentCore runtime and AgentCore reminiscence.
  • Mannequin entry granted for the fashions you intend to path to: Claude Haiku 4.5 for textual content and Claude Sonnet 4.5 for imaginative and prescient (or the equivalents accessible in your account).
  • Docker with linux/arm64 construct assist, plus the AWS Command Line Interface (AWS CLI) configured. That is wanted provided that you intend to construct and push your personal picture.
  • A Telegram bot token (from BotFather) to function the assistant’s entrance door.
  • Fundamental familiarity with agent orchestration ideas and CloudFormation.

The structure: A serverless agent on AgentCore runtime

Each element lives in a single CloudFormation template, and no construct tooling is required to launch. The next sections stroll by the load-bearing choices.

AgentCore runtime: Pay just for energetic compute

The agent lives in a container on AgentCore runtime, which makes use of consumption-based pricing. You’re billed for the compute your agent actively consumes, not for wall-clock uptime, and also you don’t pay for the time when ready for I/O corresponding to mannequin response. For a private assistant utilized in brief bursts, that’s the distinction between an roughly $1–2/month baseline and an roughly $35/month always-on Amazon Elastic Compute Cloud (Amazon EC2) occasion. These figures are estimates for mild private use as of July 2026. Consult with AgentCore pricing for present charges.

The runtime enforces a minimal container contract: hear on port 8080, and expose GET /ping for well being and POST /invocations because the agent entry level. Our container is linux/arm64, constructed multi-stage from the official OpenClaw picture plus a Python layer.

OpenClaw because the agent substrate

OpenClaw offers the agent loop, software use, and a expertise system. It runs a wrapper (server.py) that adapts it to AgentCore HTTP protocol contract:

  • On container begin, server.py launches openclaw gateway run as a subprocess and health-checks it.
  • GET /ping returns wholesome shortly, so the AgentCore readiness probe passes.
  • POST /invocations does the true work: parse the payload, retrieve reminiscence, assemble context, ahead the flip to the gateway, and persist the end result. One callout: AgentCore can thaw a frozen container whose subprocess has exited. So invocation path doesn’t assume the gateway is alive, it calls an ensure_openclaw_ready() helper that re-checks well being (and restarts the gateway if wanted) earlier than forwarding the flip.

This wrapper sample generalizes to different use instances. Any agent framework that runs as an area course of will be tailored to the AgentCore runtime the identical manner, with out modifying the framework itself.

Two fashions, routed by job

Textual content chat and picture understanding have totally different price and high quality tradeoffs, so the assistant routes them to totally different Claude fashions on Bedrock:

  • Claude Haiku 4.5 for textual content: Quick and low cost for the high-volume conversational turns that dominate day by day use.
  • Claude Sonnet 4.5 for imaginative and prescient: Stronger multimodal reasoning for the much less frequent however more durable job of diagnosing a plant from a photograph.

Textual content turns circulate by the OpenClaw gateway, which brings expertise and session state. Picture turns name the big language mannequin (LLM) from Bedrock straight from server.py, passing the picture bytes as multimodal content material blocks. We route photographs across the gateway intentionally: the in-container OpenClaw construct dropped the image_url content material components earlier than they reached Bedrock, so calling the Converse API straight from server.py makes certain the mannequin sees the precise pixels. Each paths share the identical system immediate (persona plus reminiscence), so the expertise stays constant.

The mannequin IDs are atmosphere variables (MODEL_ID, VISION_MODEL_ID), so you possibly can swap fashions per deployment with out rebuilding the picture.

Abilities because the reusable functionality unit

Capabilities are declared as expertise in a community-skills.json manifest. A deploy-time script materializes them into the container and registers them within the OpenClaw config earlier than the picture is constructed. Sprout ships with climate, reminders, and plant notes expertise on the time of publishing this publish. Swap the manifest and the identical pipeline serves a unique area. That is what makes the entire thing a reusable sample and never just one bot.

Telegram because the serverless entrance door

Telegram is a sensible channel for a private assistant because it’s webhook-based, and it retains every part serverless. It requires no consumer improvement, works on each gadget the person already owns, and helps textual content, photographs, and wealthy formatting by an easy bot API. BotFather points a bot token, which is saved in Secrets and techniques Supervisor. The deployment registers a webhook that factors Telegram on the API gateway endpoint. When the person sends a message, Telegram delivers it to the webhook Lambda perform to validate the payload and name InvokeAgentRuntime. The reply travels again by the telegram bot API.

One formatting lesson to notice: Telegram’s legacy markdown mannequin is unforgiving about unescaped characters and a single stray underscore in a mannequin response could make the entire message fail to ship. Rendering replies as HTML is dependable so the assistant converts mannequin output to Telegram-safe HTML earlier than sending.

Reminiscence: Turning disposable chats into sturdy data

The structure described up to now is a succesful, low cost, serverless agent, however by itself it nonetheless forgets you between conversations. Reminiscence is what modifications that. Think about mentioning weeks in the past that you simply backyard organically, and at present the assistant recommends a therapy and provides, by itself, that it picked the natural possibility since you don’t use artificial fertilizer. A stateless mannequin can’t do this.

AgentCore reminiscence has two layers. Brief-term reminiscence shops each dialog flip as an occasion by CreateEvent, keyed by actorId (the Telegram chat ID) and sessionId. That is the uncooked transcript. Lengthy-term reminiscence is produced asynchronously by managed extraction methods into sturdy, structured data. We configured three methods:

  • USER_PREFERENCE: specific decisions the gardener acknowledged (“I solely use natural fertilizer”).
  • SEMANTIC: inferred info (“grows Mexican petunias in a Corten metal raised mattress”).
  • SUMMARIZATION: episodic session summaries (“mentioned yellowing decrease leaves throughout a warmth wave”).

Namespaces: One backyard per gardener

Sprout information data into per-user namespaces, so no two chats ever combine:

  • sprout/{chat_id}/long_term: preferences and semantic info.
  • sprout/{chat_id}/episodic/{session_id}: session summaries.

The chat ID is the one variable phase, which makes isolation easy to purpose about and to check: every distinctive gardener maps to precisely one namespace, and no two gardeners collide.

The retrieval, meeting, and injection pipeline

On each flip, the agent retrieves the related long-term data, ranks them, and injects them into the system immediate. Here’s what occurs on each single message, inside server.py:

  • Retrieve. Name RetrieveMemoryRecords towards sprout/{chat_id}/long_term, utilizing the person’s message because the search question, capped at 50 outcomes, beneath a 3-second finances. If retrieval occasions out or errors, we degrade gracefully and reply with out reminiscence reasonably than failing.
strive:
    data = memory_client.retrieve_memory_records(
        memoryId=MEMORY_ID,
        namespace=f'sprout/{chat_id}/long_term',
        searchCriteria={
            'searchQuery': user_message,
            'topK': 50,
            'metadataFilters': []
        },
    )  # 3s timeout
besides Exception:
    data = []  # fall again to answering with out reminiscence

Snippet 1: Retrieving long-term data for the present flip (consultant. See the repo for full supply).

Assemble perform provides further customized logic. We wish the express preferences to rank forward of inferred info, order is secure inside every class, and the result’s capped earlier than injection:

def assemble(data, cap=50):
    specific = [r for r in records if r.type == 'USER_PREFERENCE']
    inferred = [r for r in records if r.type != 'USER_PREFERENCE']
    # specific beats inferred; secure order inside every class
    ordered = specific + inferred
    return ordered[:cap]

Snippet 2: The meeting step ranks specific preferences earlier than inferred info.

Metadata: Subgrouping reminiscences inside a namespace

Namespaces reply whose reminiscence a file is, however metadata solutions what it’s about. Inside sprout/{chat_id}/long_term, a semantic seek for “my petunias are wilting”, would return every part that’s shut in which means. For a gardener, meaning a fertilizer desire from March, and a fig tree pruning notice are ranked alongside data that really matter. And structured metadata helps us slender down the scope of reminiscences earlier than it reaches the immediate.

One rule shapes each resolution right here. A metadata secret is solely filterable server-side if you happen to declare it as an listed key. You may learn extra in Structured reminiscence filtering with metadata in Amazon Bedrock AgentCore Reminiscence. On this case, sprout makes use of three listed keys:

IndexedKeys:  # on the AWS::BedrockAgentCore::Reminiscence useful resource
  - Key: sort  # seperate the sorts of data
    Kind: STRING
  - Key: part  # which mattress or space it describes
    Kind: STRING
  - Key: vegetation  # what's rising there
    Kind: STRINGLIST

Every entry names a key, which should match an listed key to be filterable, and units extractionType to both STRICTLY_CONSISTENT, handed by from the occasion, or LLM_INFERRED, extracted from the dialog. For inferred keys, an extraction configuration can limit values to a hard and fast checklist. Sprout does that precisely, so each write paths would create the identical vocabulary and a filter means the identical factor no matter which half created the file.

Persisting the flip and shutting the loop

After the mannequin responds, server.py calls CreateEvent with each the person flip and the assistant flip. That new occasion feeds the extraction methods, which enrich the long-term retailer for subsequent time.

reminiscence.create_event(
    memoryId=MEMORY_ID,
    actorId=chat_id,
    sessionId=session_id,
    payload=[
        {'role': 'user', 'content': user_message},
        {'role': 'assistant', 'content': reply},
    ],
)  # feeds USER_PREFERENCE / SEMANTIC / SUMMARIZATION extraction; errors are logged, by no means deadly

Snippet 3: Persisting the flip so the extraction methods can enrich long-term reminiscence asynchronously.

Extraction is asynchronous, so a truth talked about on this session sometimes turns into retrievable in a later one. Design for that delay: short-term session occasions cowl the present dialog, and long-term data cowl every part earlier than it.

Placing it collectively: A personalised watering plan

Right here is the place the complete pipeline works end-to-end. Over just a few conversations you catalog your complete backyard, one plant at a time, in plain language. Every point out turns into an occasion. The extraction methods extract details about the plant, its location, and its solar publicity into sprout/{chat_id}/long_term. This morning the person asks a query, “Do you keep in mind the opposite vegetation in my backyard?” Retrieval pulls the data again, meeting ranks them, they usually trip into the system immediate. The assistant solutions with the person’s location, solar publicity, mattress development, soil habits, and plant stock, none of which appeared within the message itself.

Telegram chat where Sprout recalls the user’s full garden inventory, location, and sun exposure in response to a question

Determine 2: Sprout solutions a query in regards to the backyard by recalling the saved plant stock and rising circumstances

Utilizing the scheduler ability on the Amazon EventBridge → Cron path, Sprout may also flip that plan into proactive reminders (“skip the herbs, the soil continues to be damp from yesterday”) and adjusts them towards the climate ability when rain or a warmth wave is coming.

Reminiscence and imaginative and prescient additionally compound one another. When the person sends a photograph of a wilting plant, the picture goes to Claude Sonnet 4.5 whereas the system immediate nonetheless carries every part the reminiscence layer is aware of. The assistant matches the photograph to the Mexican petunias already within the person’s saved stock and diagnoses wilt stress in context reasonably than analyzing an nameless plant photograph chilly.

Telegram chat where Sprout diagnoses a wilting plant from a photo using the user’s stored Mexican petunia inventory

Determine 3: Imaginative and prescient and reminiscence working collectively. The photograph goes to the imaginative and prescient mannequin whereas the system immediate carries the person’s saved backyard context

Imaginative and prescient fashions aren’t infallible. In an earlier change with out the stock context, the identical plant was confidently recognized as a morning glory, a species with comparable trumpet-shaped purple flowers. Grounding the imaginative and prescient mannequin with the person’s personal saved stock is what turned a plausible-sounding guess into an accurate, personalised prognosis, and it’s a good illustration of why reminiscence improves accuracy and never solely tone.

Holding inference prices low with immediate caching

Injecting reminiscence into each flip makes the system immediate giant, and a naive implementation would pay for these tokens on each request. Immediate caching on Amazon Bedrock addresses this. The assistant constructions its immediate in order that the secure prefix, the persona and the assembled reminiscence block, comes first and the unstable person message comes final. Bedrock caches the processed prefix throughout requests, so repeated turns inside a dialog skip recompute of the unchanged portion. Immediate caching can cut back prices by as much as 90 % and latency by as much as 85 % for supported fashions.

The ordering rule issues greater than any single setting: put secure content material first, unstable content material final, and maintain the reminiscence block’s inside ordering deterministic (which the previous meeting perform facilitates) so the prefix really matches between requests.

Design tips to construct on AgentCore and OpenClaw

Sprout is one assistant, however the choices behind it generalize. When you’re constructing your personal assistant on this stack, the next tips are those we might carry to any area.

  • Wrap, don’t fork. Adapt your agent framework to the AgentCore container contract with a skinny HTTP wrapper reasonably than modifying the framework. The contract is small, port 8080 with /ping and /invocations, and a wrapper retains you on the framework’s improve path.
  • Design namespaces earlier than you retailer something. Reminiscence namespaces are your isolation boundary. Make the person ID the one variable phase, and select it from a channel-native ID you already belief, such because the chat ID. Multi-tenant designs get audits and deletion requests ultimately. A clear namespace scheme makes each trivial.
  • Deal with reminiscence as an enhancement, by no means a dependency. Each reminiscence operation ought to be allowed to fail gracefully. Retrieval failures ought to produce a memoryless reply with out blocking the reply. Customers forgive a forgetful flip way more readily than a failed one.
  • Route fashions by job. Use a quick, cost-effective mannequin for high-volume textual content and reserve a stronger multimodal mannequin for the turns that want it. Maintain mannequin IDs in atmosphere variables so routing modifications are configuration, not code.
  • Order prompts for the cache. Steady persona and reminiscence first, unstable person enter final, deterministic ordering all through. This one structural behavior is the place many of the inference financial savings come from.
  • Plan for extraction latency. Lengthy-term reminiscence is extracted asynchronously, so don’t promise same-session recall of recent info. Let short-term session occasions cowl the present dialog and long-term data cowl prior ones.
  • Put a finances on it from day one. A consumption-based agent is cheap till a retry loop or a chatty person makes it in any other case. An AWS Budgets alert at 80 % and one hundred pc of a month-to-month cap prices nothing and catches surprises early.
  • Maintain expertise small and single-purpose. A ability ought to do one factor a person would identify in a sentence, corresponding to examine the climate or set a reminder. Small expertise are independently testable, independently swappable, and simple for the mannequin to pick accurately. A do-everything ability forces the mannequin to guess which of its behaviors you meant.

Develop your personal

Two methods to plant it, similar backyard:

  • Single-step Launch Stack: the CloudFormation template factors at a public Amazon Elastic Container Registry (Amazon ECR) picture, so it deploys nothing however a Telegram bot token.
  • Construct your personal: The scripts/deploy.sh script validates the template, builds and pushes your personal ARM64 picture to your non-public Amazon ECR repository, deploys the stack, and registers the Telegram webhook, for a totally customizable construct.

Mild private use runs about $5–9/month as of July 2026 (roughly $2 infrastructure, $1–3 Haiku textual content, $2 Sonnet imaginative and prescient), with a built-in AWS Funds that alerts at 80 % and one hundred pc of a cap you set.

The complete supply code is accessible within the sample-agentcore-memory-openclaw GitHub repository.

Clear up

If you end up finished experimenting, tear every part right down to keep away from ongoing costs. As a result of the entire system is one CloudFormation stack, cleanup is usually a single delete:

  1. Delete the CloudFormation stack. This removes the AgentCore runtime agent, API Gateway, the Lambda capabilities, the Amazon EventBridge schedule, and the related AWS Id and Entry Administration (IAM) roles.
  2. Delete the AgentCore reminiscence retailer (and its namespaces) so no person data are retained.
  3. Delete any photographs you pushed to your non-public ECR repository, and the repository itself if it’s not wanted.
  4. Take away the AWS Funds alert if you happen to created one outdoors the stack.
  5. Revoke Telegram’s webhook (or delete the bot by BotFather), and revoke Bedrock mannequin entry if you happen to not want it.

Conclusion

The reusable core of this answer is a serverless agent on Amazon Bedrock AgentCore with a expertise system and managed reminiscence. AgentCore reminiscence removes the necessity to construct customized vector shops and extraction pipelines whereas leaving you full management over what the agent remembers and forgets, consumption-based compute plus immediate caching retains a genuinely personalised assistant at just a few {dollars} a month, and the OpenClaw expertise manifest makes the entire sample transportable throughout domains. Personalization additionally compounds: the extra a person interacts, the extra helpful the assistant turns into.

To go additional, begin with a single area corresponding to watering reminders and increase reminiscence scope incrementally, discover episodic reminiscence so the agent can reference particular previous conversations (“final time we mentioned the fig tree, you determined to carry off on fertilizer”), or fork the repository, swap in your personal persona and expertise, and develop no matter assistant you want.

To be taught extra, consult with the AgentCore documentation. The next associated posts cowl the constructing blocks in additional depth:


In regards to the authors

Thiago Verney

Thiago Verney

Thiago is a Entrance-Finish Engineer on the One MHS crew at Amazon, specializing in AI-powered interfaces utilizing React, TypeScript, and fashionable federated microfrontend structure. He builds user-focused interfaces for operations-leader in FC, drawing on prior work on the Amazon Q Developer (now Kiro) crew at AWS. He’s obsessed with innovation and crafting pragmatic options to unravel actual person issues. In his spare time, he enjoys gardening, touring, and selecting up hobbies along with his spouse in Austin, TX.

Sathya Balakrishnan

Sathya Balakrishnan

Sathya is a Pr. Cloud Architect within the Skilled Companies crew at Amazon Net Companies (AWS), specializing in knowledge and machine studying (ML) options. He works with US federal monetary purchasers. He’s obsessed with constructing pragmatic options to unravel prospects’ enterprise issues. In his spare time, he enjoys watching films and mountain climbing along with his household.

Akarsha Sehwag

Akarsha Sehwag

Akarsha is a Sr. Gen AI Information Scientist for Amazon Bedrock AgentCore GTM crew. With over 7 years of experience in AI/ML, she has constructed production-ready enterprise options throughout various buyer segments in Generative AI, Deep Studying and Pc Imaginative and prescient domains. Exterior of labor, she likes to hike, bike and play Badminton.

Tags: AgentCoreAssistantBuildingcontextawareOpenClaw
Previous Post

Agent or Workflow? A Sensible Check for Understanding When You Really Want an AI Agent

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Popular News

  • Greatest practices for Amazon SageMaker HyperPod activity governance

    Greatest practices for Amazon SageMaker HyperPod activity governance

    406 shares
    Share 162 Tweet 102
  • How Cursor Really Indexes Your Codebase

    405 shares
    Share 162 Tweet 101
  • Construct a serverless audio summarization resolution with Amazon Bedrock and Whisper

    404 shares
    Share 162 Tweet 101
  • Introducing the Amazon Bedrock AgentCore Code Interpreter

    404 shares
    Share 162 Tweet 101
  • Context Engineering — A Complete Fingers-On Tutorial with DSPy

    404 shares
    Share 162 Tweet 101

About Us

Automation Scribe is your go-to site for easy-to-understand Artificial Intelligence (AI) articles. Discover insights on AI tools, AI Scribe, and more. Stay updated with the latest advancements in AI technology. Dive into the world of automation with simplified explanations and informative content. Visit us today!

Category

  • AI Scribe
  • AI Tools
  • Artificial Intelligence

Recent Posts

  • Constructing a context-aware AI assistant on AgentCore and OpenClaw
  • Agent or Workflow? A Sensible Check for Understanding When You Really Want an AI Agent
  • How I Use AI to Be taught New Matters Quicker: An AI-Assisted Studying Framework
  • Home
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms & Conditions

© 2024 automationscribe.com. All rights reserved.

No Result
View All Result
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us

© 2024 automationscribe.com. All rights reserved.