Coding and agentic workloads are asking extra of AI fashions than ever: refactor a repository spanning a whole bunch of recordsdata, maintain a multi-hour agentic workflow with out shedding context, and purpose by way of complicated programs issues with device use at each step. Assembly these calls for with open-weight fashions has traditionally meant provisioning and working your individual inference infrastructure.
GLM 5.3 from Z.ai (Zhipu AI) is now obtainable on Amazon Bedrock. GLM 5.3, as printed on Hugging Face Hub, is a 753B-parameter mixture-of-experts mannequin optimized for coding and long-horizon agentic duties. Specifically, Z.ai has reported the mannequin reveals notable cyber safety capabilities. On Amazon Bedrock, now you can use it by way of totally managed APIs with cross-Area inference, immediate caching, and repair tiers. You don’t handle any infrastructure. Entry to GLM 5.3 on Bedrock is on the market to eligible enterprise clients.
On this submit, we present you how you can invoke GLM 5.3 on Amazon Bedrock utilizing the OpenAI-compatible APIs and cut back price and latency with immediate caching. We then put the mannequin to work in a practical agentic workflow: operating a certified safety check of your individual software with Strix, an open-source AI penetration testing agent.
What’s new in comparison with GLM 5
GLM 5 arrived on Amazon Bedrock earlier this yr. GLM 5.3 builds on the identical lineage, with a spread of necessary positive aspects:
- Stronger coding: Z.ai claims aggressive efficiency on a spread of coding benchmarks together with DeepSWE, Terminal Bench 3.0, and FrontierSWE. In addition they report a 50% enchancment over GLM 5.2 on their very own inside coding benchmark. Direct comparisons to GLM 5 weren’t reported, as a result of the magnitude of enhancements led to updating the benchmark checks themselves because the GLM 5.1 announcement.
- Emergent cyber safety capabilities: Reported benchmark efficiency on safety duties stands out, which makes the mannequin a pure match for defensive safety workflows. For instance, Z.ai measured a number one rating of 84.5 on the CyberGym benchmark at launch.
- Broader Amazon Bedrock integration: Cross-Area inference profiles, implicit and specific immediate caching, and improved function parity of the OpenAI-compatible Responses and Chat Completions APIs alongside Invoke and Converse.
Key capabilities
- Frontier coding and agentic efficiency. GLM 5.3 is designed for complicated programs engineering and long-horizon agentic duties. These embody multi-step reasoning, tool-augmented workflows, and sustained context throughout giant code bases.
- Versatile API entry. You’ll be able to invoke GLM 5.3 by way of the OpenAI-compatible Responses and Chat Completions APIs, or the Amazon Bedrock Invoke and Converse APIs.
- Immediate caching. GLM 5.3 helps implicit (computerized) immediate caching by default, and specific cache controls (advisable) on the Responses and Chat Completions APIs. For agentic workloads that resend giant system prompts or repository context each flip, caching reduces each latency and enter price.
- Cross-Area inference. GLM 5.3 is on the market by way of US cross-Area inference (
us.zai.glm-5.3) and International cross-Area inference (international.zai.glm-5.3) profiles. You ship requests to the “supply” AWS Area of your alternative, and Amazon Bedrock securely routes every request for processing. Discuss with the Amazon Bedrock Consumer Information for extra particulars. - Service tiers. Select Flex to optimize price for less-time-sensitive workloads, Precedence to prioritize latency-critical requests in return for a better value, or Commonplace for the default steadiness between value and velocity.
Stipulations
For the next utilization examples, you want:
- An AWS account with entry to Amazon Bedrock.
- AWS Id and Entry Administration (IAM) permissions to name the bottom mannequin and the goal inference profile:
bedrock:InvokeModel,bedrock:InvokeModelWithResponseStream, andbedrock:CallWithBearerToken. - (For the code-based demos) Python 3.10 or later.
- (For the non-obligatory security-testing demo solely) set up Docker and Strix with the bedrock additional.
Strive GLM 5.3 on the Amazon Bedrock console
You can begin sending prompts to GLM 5.3 on the AWS Administration Console, without having to jot down code or set up developer instruments. To get began, navigate to Amazon Bedrock after which select Check > Playground from the left sidebar menu.
From this playground interface you possibly can choose GLM 5.3 from the mannequin record and ship your first prompts by way of the chat UI, as proven within the following screenshot:
Determine 1: Chatting with GLM 5.3 on the Amazon Bedrock console
Get began with the Responses API
Programmatically, you possibly can name the mannequin by way of the bedrock-runtime endpoint. This helps each the OpenAI-compatible Responses and Chat Completions APIs, and the Amazon Bedrock Invoke and Converse APIs for GLM 5.3. For brand new purposes the OpenAI-compatible APIs are advisable as they assist a extra full set of options.
Amazon Bedrock does assist producing API keys for OpenAI-compatible integrations that require them. Nonetheless, we strongly advocate preferring short-lived credentials over long-lived API keys the place attainable.
Within the following instance, we’ll name the Responses API from Python utilizing the OpenAI Python SDK, and the aws-bedrock-token-generator library to generate short-term tokens out of your commonplace AWS Command Line Interface (AWS CLI) credentials.
- Set up the required packages.
- Save the next code as
bedrock-request.py. - Run the script, which is able to show the mannequin’s output.
Optimize inference with specific immediate caching
Lengthy-running coding and information workflows typically resend steady context throughout a number of dialog turns, reminiscent of system prompts, device definitions, or repository recordsdata.
GLM 5.3 on Amazon Bedrock helps implicit immediate caching by default, which helps cut back response latency and enter token prices for repeated calls sharing the identical preliminary immediate prefix.
With specific immediate caching mode you particularly determine the reusable immediate prefixes, which may additional enhance cache hit charge (and due to this fact latency and price financial savings) over implicit caching.
To make use of specific immediate caching with GLM 5.3, as proven within the following instance:
- Choose the express caching mode by way of
prompt_cache_optionsin your request. - Add a number of
prompt_cache_breakpointmarkers on enter content material blocks to point the top (inclusive) of reusable immediate prefixes. Every breakpoint should comprise not less than 1,024 tokens to be eligible for caching.
For extra data, consult with the immediate caching part of the Amazon Bedrock Consumer Information.
Instance agentic workload: Approved safety testing with Strix
One workload that advantages straight from GLM 5.3’s strengths is automated safety testing of your individual purposes. Strix is an open-source AI penetration testing agent that runs your code dynamically, finds vulnerabilities, and validates them with proof-of-concept checks. As of this writing, the Strix documentation makes use of GLM 5.3 as its default mannequin. You’ll be able to configure Strix to make use of GLM 5.3 on Amazon Bedrock as a substitute of a third-party inference supplier, so mannequin inference runs underneath your AWS account’s controls.
Solely check purposes you personal or have specific written permission to check. Unauthorized safety testing of programs you don’t personal is unlawful in most jurisdictions and violates the AWS Acceptable Use Coverage. On this walkthrough, the goal is OWASP Juice Store, a intentionally weak pattern software operating domestically in your machine.
If you would like totally managed, steady safety testing past operating open-source brokers your self, AWS Continuum offers on-demand penetration testing and different safety analyses as a managed service. The 2 approaches are complementary: open-source brokers like Strix provide you with developer-driven, in-the-loop, and deeply customizable testing in opposition to native builds, whereas AWS Continuum runs managed assessments at scale.
To run a certified safety check
- Begin the instance Juice Store goal software domestically.
- Configure Strix to make use of GLM 5.3 on Amazon Bedrock. Strix makes use of LiteLLM underneath the hood so (as described in their documentation for Amazon Bedrock) your AWS CLI credentials will probably be picked up mechanically. This implies no API key’s required, however you would possibly wish to set atmosphere variables like
AWS_PROFILEandAWS_REGIONto configure your connection. On the time of writing, LiteLLM doesn’t but resolvebedrock/international.zai.glm-5.3. Till that is fastened, you possibly can explicitly specify the Converse API route and the inference profile Amazon Useful resource Title (ARN) as proven within the following snippet: - Run Strix in opposition to the native goal.
- Anticipate the basis Strix agent to finish, then assessment the findings.
Strix spins up a staff of sub-agents to map the menace floor, discover a spread of potential vulnerability classes, and try to validate every discovering with a working proof of idea. This helps decrease time spent triaging false positives. A profitable run will generate a report together with severity, proof, and remediation steerage for every discovering.
The next video reveals the end-to-end journey of organising and operating Strix in opposition to the instance software, and exploring the outcomes:
Determine 2: Working an instance safety check with GLM 5.3 and Strix
Clear up
Cease the Juice Store container with Ctrl+C within the terminal the place it’s operating, or run docker ps to search out the container ID and cease it with docker cease . Amazon Bedrock inference is pay-per-token with no persistent assets, so there are not any additional prices after your requests full. When you generated an Amazon Bedrock API key for this walkthrough and now not want it, delete it on the Amazon Bedrock console.
Availability
Give GLM 5.3 a attempt on the Amazon Bedrock console, use it by way of coding assistants like OpenCode as proven in our latest submit with Kimi K3, or join your customized purposes by way of the supported APIs.
Occupied with how Amazon Bedrock can assist your staff? Join with us to start out the dialog.
In regards to the authors

