Automationscribe.com
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us
No Result
View All Result
Automation Scribe
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us
No Result
View All Result
Automationscribe.com
No Result
View All Result

Introducing GLM 5.3 on Amazon Bedrock

admin by admin
October 6, 2026
in Artificial Intelligence
0
Introducing GLM 5.3 on Amazon Bedrock
399
SHARES
2.3k
VIEWS
Share on FacebookShare on Twitter


Coding and agentic workloads are asking extra of AI fashions than ever: refactor a repository spanning a whole bunch of recordsdata, maintain a multi-hour agentic workflow with out shedding context, and purpose by way of complicated programs issues with device use at each step. Assembly these calls for with open-weight fashions has traditionally meant provisioning and working your individual inference infrastructure.

GLM 5.3 from Z.ai (Zhipu AI) is now obtainable on Amazon Bedrock. GLM 5.3, as printed on Hugging Face Hub, is a 753B-parameter mixture-of-experts mannequin optimized for coding and long-horizon agentic duties. Specifically, Z.ai has reported the mannequin reveals notable cyber safety capabilities. On Amazon Bedrock, now you can use it by way of totally managed APIs with cross-Area inference, immediate caching, and repair tiers. You don’t handle any infrastructure. Entry to GLM 5.3 on Bedrock is on the market to eligible enterprise clients.

On this submit, we present you how you can invoke GLM 5.3 on Amazon Bedrock utilizing the OpenAI-compatible APIs and cut back price and latency with immediate caching. We then put the mannequin to work in a practical agentic workflow: operating a certified safety check of your individual software with Strix, an open-source AI penetration testing agent.

What’s new in comparison with GLM 5

GLM 5 arrived on Amazon Bedrock earlier this yr. GLM 5.3 builds on the identical lineage, with a spread of necessary positive aspects:

  • Stronger coding: Z.ai claims aggressive efficiency on a spread of coding benchmarks together with DeepSWE, Terminal Bench 3.0, and FrontierSWE. In addition they report a 50% enchancment over GLM 5.2 on their very own inside coding benchmark. Direct comparisons to GLM 5 weren’t reported, as a result of the magnitude of enhancements led to updating the benchmark checks themselves because the GLM 5.1 announcement.
  • Emergent cyber safety capabilities: Reported benchmark efficiency on safety duties stands out, which makes the mannequin a pure match for defensive safety workflows. For instance, Z.ai measured a number one rating of 84.5 on the CyberGym benchmark at launch.
  • Broader Amazon Bedrock integration: Cross-Area inference profiles, implicit and specific immediate caching, and improved function parity of the OpenAI-compatible Responses and Chat Completions APIs alongside Invoke and Converse.

Key capabilities

  • Frontier coding and agentic efficiency. GLM 5.3 is designed for complicated programs engineering and long-horizon agentic duties. These embody multi-step reasoning, tool-augmented workflows, and sustained context throughout giant code bases.
  • Versatile API entry. You’ll be able to invoke GLM 5.3 by way of the OpenAI-compatible Responses and Chat Completions APIs, or the Amazon Bedrock Invoke and Converse APIs.
  • Immediate caching. GLM 5.3 helps implicit (computerized) immediate caching by default, and specific cache controls (advisable) on the Responses and Chat Completions APIs. For agentic workloads that resend giant system prompts or repository context each flip, caching reduces each latency and enter price.
  • Cross-Area inference. GLM 5.3 is on the market by way of US cross-Area inference (us.zai.glm-5.3) and International cross-Area inference (international.zai.glm-5.3) profiles. You ship requests to the “supply” AWS Area of your alternative, and Amazon Bedrock securely routes every request for processing. Discuss with the Amazon Bedrock Consumer Information for extra particulars.
  • Service tiers. Select Flex to optimize price for less-time-sensitive workloads, Precedence to prioritize latency-critical requests in return for a better value, or Commonplace for the default steadiness between value and velocity.

Stipulations

For the next utilization examples, you want:

  1. An AWS account with entry to Amazon Bedrock.
  2. AWS Id and Entry Administration (IAM) permissions to name the bottom mannequin and the goal inference profile: bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, and bedrock:CallWithBearerToken.
  3. (For the code-based demos) Python 3.10 or later.
  4. (For the non-obligatory security-testing demo solely) set up Docker and Strix with the bedrock additional.

Strive GLM 5.3 on the Amazon Bedrock console

You can begin sending prompts to GLM 5.3 on the AWS Administration Console, without having to jot down code or set up developer instruments. To get began, navigate to Amazon Bedrock after which select Check > Playground from the left sidebar menu.

From this playground interface you possibly can choose GLM 5.3 from the mannequin record and ship your first prompts by way of the chat UI, as proven within the following screenshot:

[Amazon Bedrock Playground screenshot showing chat interface with GLM 5.3 model explaining symmetric vs asymmetric encryption. Response includes detailed comparison with checkmarks and X marks highlighting key differences, examples like AES and RSA, a]

Determine 1: Chatting with GLM 5.3 on the Amazon Bedrock console

Get began with the Responses API

Programmatically, you possibly can name the mannequin by way of the bedrock-runtime endpoint. This helps each the OpenAI-compatible Responses and Chat Completions APIs, and the Amazon Bedrock Invoke and Converse APIs for GLM 5.3. For brand new purposes the OpenAI-compatible APIs are advisable as they assist a extra full set of options.

Amazon Bedrock does assist producing API keys for OpenAI-compatible integrations that require them. Nonetheless, we strongly advocate preferring short-lived credentials over long-lived API keys the place attainable.

Within the following instance, we’ll name the Responses API from Python utilizing the OpenAI Python SDK, and the aws-bedrock-token-generator library to generate short-term tokens out of your commonplace AWS Command Line Interface (AWS CLI) credentials.

  1. Set up the required packages.
    pip set up -U openai aws-bedrock-token-generator

  2. Save the next code as bedrock-request.py.
    from aws_bedrock_token_generator import provide_token
    from openai import OpenAI
    
    area = "us-west-2"  # Your supply AWS Area
    
    consumer = OpenAI(
        api_key=provide_token(area=area),
        base_url=f"https://bedrock-runtime.{area}.amazonaws.com/openai/v1",
    )
    
    resp = consumer.responses.create(
        enter="Refactor this Python perform to be iterative as a substitute of recursive: ...",
        mannequin="international.zai.glm-5.3",
    )
    
    print(resp.output_text)

  3. Run the script, which is able to show the mannequin’s output.
    python bedrock-request.py

Optimize inference with specific immediate caching

Lengthy-running coding and information workflows typically resend steady context throughout a number of dialog turns, reminiscent of system prompts, device definitions, or repository recordsdata.

GLM 5.3 on Amazon Bedrock helps implicit immediate caching by default, which helps cut back response latency and enter token prices for repeated calls sharing the identical preliminary immediate prefix.

With specific immediate caching mode you particularly determine the reusable immediate prefixes, which may additional enhance cache hit charge (and due to this fact latency and price financial savings) over implicit caching.

To make use of specific immediate caching with GLM 5.3, as proven within the following instance:

  1. Choose the express caching mode by way of prompt_cache_options in your request.
  2. Add a number of prompt_cache_breakpoint markers on enter content material blocks to point the top (inclusive) of reusable immediate prefixes. Every breakpoint should comprise not less than 1,024 tokens to be eligible for caching.
resp = consumer.responses.create(
    mannequin="international.zai.glm-5.3",
    # Allow specific caching mode:
    extra_body={"prompt_cache_options": {"mode": "specific"}},
    enter=[
        {
            "type": "message",
            "role": "system",
            "content": [
                {
                    "type": "input_text",
                    "text": SYSTEM_PROMPT,
                    # A long, static system prompt is a great target for caching:
                    "prompt_cache_breakpoint": {"mode": "explicit"},
                },
            ]
        },
        {
            "kind": "message",
            "function": "person",
            "content material": [
                {
                    "type": "input_text",
                    "text": USER_INPUT,
                    # Multiple breakpoints can also be defined, for layered cache:
                    "prompt_cache_breakpoint": {"mode": "explicit"},
                },
            ],
        },
    ],
)

if resp.utilization.input_tokens_details.cached_tokens:
    print("Hit cache!")

For extra data, consult with the immediate caching part of the Amazon Bedrock Consumer Information.

Instance agentic workload: Approved safety testing with Strix

One workload that advantages straight from GLM 5.3’s strengths is automated safety testing of your individual purposes. Strix is an open-source AI penetration testing agent that runs your code dynamically, finds vulnerabilities, and validates them with proof-of-concept checks. As of this writing, the Strix documentation makes use of GLM 5.3 as its default mannequin. You’ll be able to configure Strix to make use of GLM 5.3 on Amazon Bedrock as a substitute of a third-party inference supplier, so mannequin inference runs underneath your AWS account’s controls.

Solely check purposes you personal or have specific written permission to check. Unauthorized safety testing of programs you don’t personal is unlawful in most jurisdictions and violates the AWS Acceptable Use Coverage. On this walkthrough, the goal is OWASP Juice Store, a intentionally weak pattern software operating domestically in your machine.

If you would like totally managed, steady safety testing past operating open-source brokers your self, AWS Continuum offers on-demand penetration testing and different safety analyses as a managed service. The 2 approaches are complementary: open-source brokers like Strix provide you with developer-driven, in-the-loop, and deeply customizable testing in opposition to native builds, whereas AWS Continuum runs managed assessments at scale.

To run a certified safety check

  1. Begin the instance Juice Store goal software domestically.
    docker run --rm -p 3000:3000 bkimminich/juice-shop

  2. Configure Strix to make use of GLM 5.3 on Amazon Bedrock. Strix makes use of LiteLLM underneath the hood so (as described in their documentation for Amazon Bedrock) your AWS CLI credentials will probably be picked up mechanically. This implies no API key’s required, however you would possibly wish to set atmosphere variables like AWS_PROFILE and AWS_REGION to configure your connection. On the time of writing, LiteLLM doesn’t but resolve bedrock/international.zai.glm-5.3. Till that is fastened, you possibly can explicitly specify the Converse API route and the inference profile Amazon Useful resource Title (ARN) as proven within the following snippet:
    # Fill within the REGION and ACCOUNT_ID placeholders beneath earlier than operating!
    export STRIX_LLM="bedrock/converse/arn:aws:bedrock:{AWS_REGION}:{AWS_ACCOUNT_ID}:inference-profile/international.zai.glm-5.3"

  3. Run Strix in opposition to the native goal.
    strix --target http://localhost:3000

  4. Anticipate the basis Strix agent to finish, then assessment the findings.

Strix spins up a staff of sub-agents to map the menace floor, discover a spread of potential vulnerability classes, and try to validate every discovering with a working proof of idea. This helps decrease time spent triaging false positives. A profitable run will generate a report together with severity, proof, and remediation steerage for every discovering.

The next video reveals the end-to-end journey of organising and operating Strix in opposition to the instance software, and exploring the outcomes:

Determine 2: Working an instance safety check with GLM 5.3 and Strix

Clear up

Cease the Juice Store container with Ctrl+C within the terminal the place it’s operating, or run docker ps to search out the container ID and cease it with docker cease . Amazon Bedrock inference is pay-per-token with no persistent assets, so there are not any additional prices after your requests full. When you generated an Amazon Bedrock API key for this walkthrough and now not want it, delete it on the Amazon Bedrock console.

Availability

Give GLM 5.3 a attempt on the Amazon Bedrock console, use it by way of coding assistants like OpenCode as proven in our latest submit with Kimi K3, or join your customized purposes by way of the supported APIs.

Occupied with how Amazon Bedrock can assist your staff? Join with us to start out the dialog.


In regards to the authors

Alex Thewsey

Alex Thewsey

Alex is an AI Specialist Options Architect at AWS, primarily based in Singapore. He focuses on how open supply applied sciences and open weight fashions may also help clients around the globe to construct modern AI options and deal with AI governance challenges.

Tags: AmazonBedrockGLMIntroducing
Previous Post

AI Agent Observability: Logging, Tracing, and Debugging Defined

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Popular News

  • Greatest practices for Amazon SageMaker HyperPod activity governance

    Greatest practices for Amazon SageMaker HyperPod activity governance

    406 shares
    Share 162 Tweet 102
  • How Cursor Really Indexes Your Codebase

    405 shares
    Share 162 Tweet 101
  • Construct a serverless audio summarization resolution with Amazon Bedrock and Whisper

    404 shares
    Share 162 Tweet 101
  • Context Engineering — A Complete Fingers-On Tutorial with DSPy

    404 shares
    Share 162 Tweet 101
  • Introducing the Amazon Bedrock AgentCore Code Interpreter

    404 shares
    Share 162 Tweet 101

About Us

Automation Scribe is your go-to site for easy-to-understand Artificial Intelligence (AI) articles. Discover insights on AI tools, AI Scribe, and more. Stay updated with the latest advancements in AI technology. Dive into the world of automation with simplified explanations and informative content. Visit us today!

Category

  • AI Scribe
  • AI Tools
  • Artificial Intelligence

Recent Posts

  • Introducing GLM 5.3 on Amazon Bedrock
  • AI Agent Observability: Logging, Tracing, and Debugging Defined
  • Laptop Imaginative and prescient: SIFT algorithm (Scale Invariant Function Rework)
  • Home
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms & Conditions

© 2024 automationscribe.com. All rights reserved.

No Result
View All Result
  • Home
  • AI Scribe
  • AI Tools
  • Artificial Intelligence
  • Contact Us

© 2024 automationscribe.com. All rights reserved.