From AWS Machine Learning Blog · 40 stories
NVIDIA Launches Revenue-Sharing Model for AI Infrastructure and Agent Toolkit
NVIDIA has introduced a revenue-sharing model for AI cloud partners to access its infrastructure more affordably, enabling startups to pay a percentage of revenue in addition to hardware costs. Additionally, NVIDIA released an Agent Toolkit to facilitate the creation of specialized AI systems within business workflows. These initiatives aim to expand NVIDIA's AI technology reach and revenue sources.
Strategic Frameworks and Systems Vital for Successful AI Integration in Enterprises
AI's integration in enterprises is moving beyond model development to focus on creating robust systems for execution and governance. This shift highlights the importance of developing adaptable frameworks to support AI's role across various functions such as finance, HR, and operations. It reflects a broader industry trend where the focus is on building the necessary infrastructure to ensure AI's ongoing, safe, and productive incorporation into real-world workflows, addressing the current challenges and limitations.
Meta Introduces Hybrid Asset Classification for Privacy-Aware Infrastructure
Meta has unveiled a hybrid asset classification strategy using large language models (LLMs) to handle ambiguous data in privacy-aware infrastructure while maintaining deterministic rules for enforcement. This method addresses the complexities of AI-native products with varied data inputs, ensuring compliance and effective data governance. It is a response to the challenges posed by the increasing speed and scale of AI innovations, and the approach aims to better manage privacy controls for evolving AI products.
Amazon Bedrock Enhances AI Capabilities with Security and Operational Features
Amazon Bedrock has introduced several updates to improve the security and operational management of AI applications, emphasizing capabilities for multi-tenant AI, data retention policies, and compliance with US government standards. Key features include resource-based policies, managed entitlements for model subscriptions, zero data retention enforcement, and AI model support in AWS GovCloud. These advancements aim to streamline AI adoption across diverse sectors while maintaining security and governance standards.
U.S. Greenlights Public Rollout of OpenAI's GPT-5.6 Amid Regulatory Controls
The U.S. government has approved OpenAI's GPT-5.6 models for public release on July 9, ending a period of limited access due to regulatory scrutiny. The launch of these models, including Sol, Terra, and Luna, follows compliance with federal cybersecurity reviews intended to manage AI model rollouts. This episode highlights the tension between advancing AI capabilities and the increasing regulatory oversight.
New AI Models for Long-Horizon Coding Tasks Introduced
Several AI models aimed at long-horizon tasks in coding and robotics have been released. GLM-5.2 by Hugging Face extends support for coding-agent scenarios with a 1 million token context. Cognition's SWE-1.7 enhances long-horizon asynchronous tasks with reinforcement learning. Xiaomi-Robotics-1 combines vast pre-training data for improved robotics capabilities. These releases highlight advances in scaling and reasoning capabilities.
SpaceXAI Releases Grok Bot AI Agent and Grok 4.6 Model, Completes Cursor Acquisition
SpaceXAI, which recently completed its acquisition of AI coding company Cursor, has launched Grok Bot, an AI agent for Mac, iOS, Windows, and Linux, designed to automate tasks across applications. Concurrently, the company released Grok 4.6, an updated AI model that scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol and offering competitive pricing for long-running agents, coding, and knowledge work.
Challenges in AI Token Costs and Efficiency Revealed
AI model pricing based on tokens has been criticized for being misleading due to varying tokenization methods. DeepSeek's price cut on its V4-Pro model exemplifies the complexity as lower token rates don't guarantee cost savings. Researchers highlight solutions like AI harnesses that optimize token usage, offering cost-effective alternatives.
OpenAI's GPT-6 Astra Achieves High Scores on ARC-AGI-3 Benchmark
OpenAI has released GPT-6 Astra, its latest AI model, which achieved a 99.9% score on the ARC-AGI-3 benchmark using a provider adapter harness. This marks a significant improvement over its predecessor, GPT-5.6 Sol, which scored 7.8%, and demonstrates the model's ability to navigate unfamiliar interactive environments and create symbolic world models.
Claude Mythos 5 AI Cybersecurity Capabilities Expanded, $35M Fund for Open-Source Security
Claude Mythos 5, an AI model for cybersecurity, is now available in Claude Security and will integrate into partner tools. The company also launched a $35 million fund to support open-source software security and plans to expand its Cyber Verification Program. These actions aim to broaden access to advanced AI for defensive cybersecurity while maintaining safeguards against misuse.
Amazon Quick Enhances Efficiency for Finance, Sales, and Supply Chain Management
Amazon Quick, a generative AI assistant, enhances efficiency for AWS Finance, sales teams, Tradeshift, and supply chain operations. It streamlines data preparation, administrative tasks, and analytics processes, allowing organizations to improve productivity and response times across various sectors.
Anthropic Releases Claude Opus 5, Offering Near Fable 5 Performance at Half the Cost
Anthropic has released Claude Opus 5, a new AI model that approaches the intelligence of its flagship Fable 5 model but at half the price. Opus 5 is now the default model for Claude Max subscribers and the strongest available for Claude Pro users, offering improved performance in coding and knowledge work while maintaining the token cost of its predecessor, Opus 4.8.
Amazon SageMaker AI Integrates with MLflow for Enhanced Model Deployment and Monitoring
Amazon SageMaker AI has integrated with MLflow to provide a unified tracking interface for AI inference, benchmark, and model monitoring. This supports both generative AI models and discriminative ML models, simplifying model deployment and improving ongoing accuracy by detecting data drift. The integration is designed to reduce setup time and enhance workflow reproducibility in machine learning applications.
Alibaba Releases Qwen3.8-Max and Qwen3.8-27B AI Models, Including Open Weights
Alibaba has released its Qwen3.8-Max and Qwen3.8-27B AI models. Qwen3.8-Max is a multimodal model with 2.4 trillion parameters and a 1 million token context window, while Qwen3.8-27B is a 27-billion-parameter version with native vision-language understanding. The company plans to release open weights for both models, making Qwen-Max-class capabilities available to the open-source community.
Model Context Protocol (MCP) 2026-07-28 Specification Released, Adopting Stateless Core
The Model Context Protocol (MCP) has released its 2026-07-28 specification, transitioning from a bidirectional stateful protocol to a request/response stateless core. This update, the largest revision since its launch, aims to improve reliability and scalability for MCP servers, addressing a highly requested developer feature. Major SDKs, including TypeScript, Python, and C#, have been updated to support the new specification.
SpaceXAI Launches Cost-Effective AI Model Grok 4.5, Challenging Rivals
SpaceXAI released Grok 4.5, a new AI model built in collaboration with Cursor, focusing on coding and engineering tasks. The model offers improved efficiency and lower costs, presenting a competitive alternative to existing AI models. Grok 4.5's release introduces a pricing strategy that undercuts rivals, influencing the AI market dynamics among enterprise users and developers.
Anthropic Releases Economical AI Model Claude Sonnet 5 for Enhanced Agentic Tasks
Anthropic has unveiled Claude Sonnet 5, an advanced AI model designed for agentic capabilities, available on AWS. This model provides substantial improvements in planning, tool use, and coding over previous versions at a lower cost, aiming to compete with top-tier AI models.
Inscribe Uses Amazon Bedrock for Rapid Document Fraud Detection
Inscribe utilizes Amazon Bedrock to enhance its document fraud detection, reducing verification time to under 90 seconds. This advancement addresses the surge in AI-generated fraud and the pressing need for financial institutions to maintain accuracy amidst increasing application volumes.
Kimi K3: New 2.8 Trillion-Parameter Open AI Model with Vision Capabilities Released
Kimi has introduced Kimi K3, a 2.8 trillion-parameter AI model with built-in vision capabilities and a 1-million-token context window. This open-source model, while not as powerful as proprietary models, shows competitive performance across various benchmarks, including a strong showing against Fable 5. Kimi K3 offers new potential for AI applications but faces cost and speed challenges.
Anthropic Launches Claude Apps Gateway for AWS and Google Cloud
Anthropic has launched the Claude apps gateway for both AWS and Google Cloud. This tool provides enterprises with centralized management of Claude Code, enhancing control over access, costs, and policy adherence. By streamlining developer credential management and spend tracking, the gateway eases administrative burdens for organizations operating at scale.
Claude Sonnet 5.5 now available on Amazon Bedrock and Claude Platform on AWS
Claude Sonnet 5.5, an updated AI model, is now available on Amazon Bedrock and Claude Platform on AWS. This version offers improved efficiency for coding and knowledge work at a lower cost and faster speed, making it suitable for continuous or large-scale AI workloads within AWS infrastructure.
Amazon SageMaker HyperPod integrates new Ray capabilities for distributed ML workloads
Amazon SageMaker HyperPod now includes new capabilities for Ray, an open-source framework for scaling distributed Python workloads. This integration simplifies the management of Ray clusters on HyperPod's infrastructure, providing built-in fault tolerance and improved observability for large-scale machine learning.
AWS SageMaker AI now supports WhisperX for speaker-labeled transcription
AWS SageMaker AI now supports WhisperX, an extension of OpenAI's Whisper model, for speaker-labeled transcription. This integration provides precise per-word timestamps and speaker diarization, addressing limitations of generic speech-to-text solutions for various enterprise workloads.
Amazon Bedrock Introduces Prompt Caching to Reduce Costs and Latency
Amazon Bedrock now offers prompt caching, which can reduce input token costs by up to 90% and lower time-to-first-token (TTFT) when repeatedly sending the same context to foundation models. This feature allows Bedrock to store and reuse partially processed input, avoiding redundant computation for subsequent requests with matching prefixes.
TwelveLabs Marengo Embed 3.0 now available in Amazon Bedrock Knowledge Bases
TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model within Amazon Bedrock Knowledge Bases. This integration allows for natural language search across video, audio, and image content, simplifying the process of building semantic search capabilities for multimodal assets.
AWS introduces Ray Serve Deep Learning Containers for TorchServe inference workloads
AWS has launched Ray Serve Deep Learning Containers (DLCs) to support model inference workloads, particularly for users migrating from the unmaintained TorchServe. These DLCs provide pre-built, optimized Docker images with a complete inference stack, including PyTorch and Ray Serve, maintained by AWS. This offering addresses the operational burden of managing dependencies and security for deep learning model serving.
NVIDIA MPS and Triton on Amazon EC2 Reduce ASR Inference Costs by 75%
AWS, NVIDIA, and Heidi Health demonstrated a 75% reduction in ASR inference costs on Amazon EC2 by using NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server. This setup allows for more efficient GPU utilization, reducing the number of required GPU instances from 16 to 4 for Heidi Health's clinical consultation processing. The improved efficiency addresses the challenge of low GPU utilization during ASR inference while maintaining sub-second latency.
Amazon SageMaker HyperPod introduces tiered KV cache with Curvine for LLM inference
Amazon SageMaker HyperPod now supports a tiered KV cache architecture using Curvine, a distributed cache filesystem, to extend KV cache beyond GPU and CPU memory into a shared NVMe pool. This development allows for KV cache reuse across replicas, improving time-to-first-token (TTFT) and reducing infrastructure costs for large language model (LLM) inference.
Amazon Bedrock AgentCore Identity Adds Private Key JWT Authentication
Amazon Bedrock AgentCore Identity now supports Private Key JWT client authentication for agents, allowing them to authenticate to identity providers using signed JSON Web Tokens instead of shared secrets. This enhancement improves security by keeping private keys within AWS Key Management Service (KMS) and enabling more secure machine-to-machine authentication flows for AI agents.
Amazon Nova introduces Self-Distilled Reasoning for enhanced model fine-tuning
Amazon's Nova 2 models now utilize Self-Distilled Reasoning (SDR) to improve fine-tuning. SDR improves model performance on datasets lacking reasoning traces, addressing issues like catastrophic forgetting during training.
Custom OS Installation Enabled for AWS DeepRacer Devices
AWS has launched a developer bootloader for AWS DeepRacer devices allowing for custom OS installations. This update enables developers to utilize modern Linux distributions and tailor the devices for enhanced functionality and educational use.
Amazon Quick adds Mobile Layout for Free Form dashboards
Amazon Quick now offers a Mobile Layout feature for Free Form dashboards, allowing automatic rendering in a single-column view for mobile devices. This change facilitates easier access to data on phones and tablets, enhancing user experience without requiring any modifications from dashboard authors.
Smartsheet Develops Remote MCP Server on AWS for AI Integration
Smartsheet has built a remote Model Context Protocol (MCP) server on AWS to enable direct AI access to its data. This development allows AI agents to interact with Smartsheet's capabilities, enhancing workflow efficiency for enterprise users.
Built Technologies launches AI document intelligence solution on AWS for real estate finance
Built Technologies has launched an AI-powered document intelligence solution on AWS to enhance processing in real estate finance. This solution aims to automate complex and manual document processing, significantly reducing workflow times and supporting various document types throughout the real estate lifecycle.
Amazon introduces custom CloudWatch dashboards for SageMaker Pipelines monitoring
Amazon has released a new solution for monitoring SageMaker Pipelines across multiple AWS accounts and Regions using customized CloudWatch dashboards. This centralization helps organizations manage their machine learning operations more efficiently and reduces the operational overhead associated with monitoring distributed environments.
Amazon Nova Act Introduces Scalable UX Testing Approach for User Flow Analysis
Amazon launched Nova Act, a multimodal foundation model designed to enhance User Experience (UX) testing by utilizing visual understanding for intelligent navigation. This approach allows for more comprehensive testing of user workflows while adapting to interface changes, addressing the limitations of traditional automated testing methods.
Flo Health Implements AI Content Review System with Amazon Bedrock
Flo Health's engineering team has developed a production-grade medical content review system using Amazon Bedrock, resulting in a 60% reduction in review time and tripled content throughput. This advancement addresses the challenges of scaling medical content production while maintaining rigorous accuracy standards without the need for expanding the medical team.
ScienceSoft launches HIPAA-compliant AI voice scheduler using AWS
ScienceSoft has developed a HIPAA-compliant AI voice scheduler leveraging Amazon technologies to streamline healthcare appointment scheduling. This solution addresses inefficiencies in traditional manual workflows, potentially transforming patient scheduling in healthcare settings.
Henry Schein One implements AI for real-time dental image verification
Henry Schein One launched Image Verify, an AI system using Amazon SageMaker, to assess dental X-ray quality in real-time. This innovation addresses a significant bottleneck in dental claims processing, minimizing delays and improving patient care as the system scales to 40,000 locations globally.
Amazon Quick Automate Introduces Native Case Management for Workflows
Amazon Quick Automate now includes native case management, enhancing the tracking and management of workflows across AI agents. This feature enables organizations to monitor work item progress, handle exceptions, and allow human intervention when necessary, facilitating efficient enterprise-scale automation.