From AWS Machine Learning Blog · 40 stories
NVIDIA Launches Revenue-Sharing Model for AI Infrastructure and Agent Toolkit
NVIDIA has introduced a revenue-sharing model for AI cloud partners to access its infrastructure more affordably, enabling startups to pay a percentage of revenue in addition to hardware costs. Additionally, NVIDIA released an Agent Toolkit to facilitate the creation of specialized AI systems within business workflows. These initiatives aim to expand NVIDIA's AI technology reach and revenue sources.
Meta Introduces Hybrid Asset Classification for Privacy-Aware Infrastructure
Meta has unveiled a hybrid asset classification strategy using large language models (LLMs) to handle ambiguous data in privacy-aware infrastructure while maintaining deterministic rules for enforcement. This method addresses the complexities of AI-native products with varied data inputs, ensuring compliance and effective data governance. It is a response to the challenges posed by the increasing speed and scale of AI innovations, and the approach aims to better manage privacy controls for evolving AI products.
Strategic Frameworks and Systems Vital for Successful AI Integration in Enterprises
AI's integration in enterprises is moving beyond model development to focus on creating robust systems for execution and governance. This shift highlights the importance of developing adaptable frameworks to support AI's role across various functions such as finance, HR, and operations. It reflects a broader industry trend where the focus is on building the necessary infrastructure to ensure AI's ongoing, safe, and productive incorporation into real-world workflows, addressing the current challenges and limitations.
Amazon Bedrock Enhances AI Capabilities with Security and Operational Features
Amazon Bedrock has introduced several updates to improve the security and operational management of AI applications, emphasizing capabilities for multi-tenant AI, data retention policies, and compliance with US government standards. Key features include resource-based policies, managed entitlements for model subscriptions, zero data retention enforcement, and AI model support in AWS GovCloud. These advancements aim to streamline AI adoption across diverse sectors while maintaining security and governance standards.
Model Context Protocol (MCP) 2026-07-28 Specification Released, Adopting Stateless Core
The Model Context Protocol (MCP) has released its 2026-07-28 specification, transitioning from a bidirectional stateful protocol to a request/response stateless core. This update, the largest revision since its launch, aims to improve reliability and scalability for MCP servers, addressing a highly requested developer feature. Major SDKs, including TypeScript, Python, and C#, have been updated to support the new specification.
U.S. Greenlights Public Rollout of OpenAI's GPT-5.6 Amid Regulatory Controls
The U.S. government has approved OpenAI's GPT-5.6 models for public release on July 9, ending a period of limited access due to regulatory scrutiny. The launch of these models, including Sol, Terra, and Luna, follows compliance with federal cybersecurity reviews intended to manage AI model rollouts. This episode highlights the tension between advancing AI capabilities and the increasing regulatory oversight.
Anthropic Releases Claude Opus 5, Offering Near Fable 5 Performance at Half the Cost
Anthropic has released Claude Opus 5, a new AI model that approaches the intelligence of its flagship Fable 5 model but at half the price. Opus 5 is now the default model for Claude Max subscribers and the strongest available for Claude Pro users, offering improved performance in coding and knowledge work while maintaining the token cost of its predecessor, Opus 4.8.
SpaceXAI Launches Cost-Effective AI Model Grok 4.5, Challenging Rivals
SpaceXAI released Grok 4.5, a new AI model built in collaboration with Cursor, focusing on coding and engineering tasks. The model offers improved efficiency and lower costs, presenting a competitive alternative to existing AI models. Grok 4.5's release introduces a pricing strategy that undercuts rivals, influencing the AI market dynamics among enterprise users and developers.
Anthropic Releases Economical AI Model Claude Sonnet 5 for Enhanced Agentic Tasks
Anthropic has unveiled Claude Sonnet 5, an advanced AI model designed for agentic capabilities, available on AWS. This model provides substantial improvements in planning, tool use, and coding over previous versions at a lower cost, aiming to compete with top-tier AI models.
Inscribe Uses Amazon Bedrock for Rapid Document Fraud Detection
Inscribe utilizes Amazon Bedrock to enhance its document fraud detection, reducing verification time to under 90 seconds. This advancement addresses the surge in AI-generated fraud and the pressing need for financial institutions to maintain accuracy amidst increasing application volumes.
Amazon Quick Enhances Efficiency for Finance, Sales, and Supply Chain Management
Amazon Quick, a generative AI assistant, enhances efficiency for AWS Finance, sales teams, Tradeshift, and supply chain operations. It streamlines data preparation, administrative tasks, and analytics processes, allowing organizations to improve productivity and response times across various sectors.
Kimi K3: New 2.8 Trillion-Parameter Open AI Model with Vision Capabilities Released
Kimi has introduced Kimi K3, a 2.8 trillion-parameter AI model with built-in vision capabilities and a 1-million-token context window. This open-source model, while not as powerful as proprietary models, shows competitive performance across various benchmarks, including a strong showing against Fable 5. Kimi K3 offers new potential for AI applications but faces cost and speed challenges.
Anthropic Launches Claude Apps Gateway for AWS and Google Cloud
Anthropic has launched the Claude apps gateway for both AWS and Google Cloud. This tool provides enterprises with centralized management of Claude Code, enhancing control over access, costs, and policy adherence. By streamlining developer credential management and spend tracking, the gateway eases administrative burdens for organizations operating at scale.
Amazon SageMaker HyperPod introduces tiered KV cache with Curvine for LLM inference
Amazon SageMaker HyperPod now supports a tiered KV cache architecture using Curvine, a distributed cache filesystem, to extend KV cache beyond GPU and CPU memory into a shared NVMe pool. This development allows for KV cache reuse across replicas, improving time-to-first-token (TTFT) and reducing infrastructure costs for large language model (LLM) inference.
Amazon SageMaker AI Integrates with MLflow for Enhanced Model Deployment and Monitoring
Amazon SageMaker AI has integrated with MLflow to provide a unified tracking interface for AI inference, benchmark, and model monitoring. This supports both generative AI models and discriminative ML models, simplifying model deployment and improving ongoing accuracy by detecting data drift. The integration is designed to reduce setup time and enhance workflow reproducibility in machine learning applications.
Amazon Bedrock AgentCore Identity Adds Private Key JWT Authentication
Amazon Bedrock AgentCore Identity now supports Private Key JWT client authentication for agents, allowing them to authenticate to identity providers using signed JSON Web Tokens instead of shared secrets. This enhancement improves security by keeping private keys within AWS Key Management Service (KMS) and enabling more secure machine-to-machine authentication flows for AI agents.
Amazon Nova introduces Self-Distilled Reasoning for enhanced model fine-tuning
Amazon's Nova 2 models now utilize Self-Distilled Reasoning (SDR) to improve fine-tuning. SDR improves model performance on datasets lacking reasoning traces, addressing issues like catastrophic forgetting during training.
Custom OS Installation Enabled for AWS DeepRacer Devices
AWS has launched a developer bootloader for AWS DeepRacer devices allowing for custom OS installations. This update enables developers to utilize modern Linux distributions and tailor the devices for enhanced functionality and educational use.
Amazon Quick adds Mobile Layout for Free Form dashboards
Amazon Quick now offers a Mobile Layout feature for Free Form dashboards, allowing automatic rendering in a single-column view for mobile devices. This change facilitates easier access to data on phones and tablets, enhancing user experience without requiring any modifications from dashboard authors.
Smartsheet Develops Remote MCP Server on AWS for AI Integration
Smartsheet has built a remote Model Context Protocol (MCP) server on AWS to enable direct AI access to its data. This development allows AI agents to interact with Smartsheet's capabilities, enhancing workflow efficiency for enterprise users.
Built Technologies launches AI document intelligence solution on AWS for real estate finance
Built Technologies has launched an AI-powered document intelligence solution on AWS to enhance processing in real estate finance. This solution aims to automate complex and manual document processing, significantly reducing workflow times and supporting various document types throughout the real estate lifecycle.
Amazon introduces custom CloudWatch dashboards for SageMaker Pipelines monitoring
Amazon has released a new solution for monitoring SageMaker Pipelines across multiple AWS accounts and Regions using customized CloudWatch dashboards. This centralization helps organizations manage their machine learning operations more efficiently and reduces the operational overhead associated with monitoring distributed environments.
Amazon Nova Act Introduces Scalable UX Testing Approach for User Flow Analysis
Amazon launched Nova Act, a multimodal foundation model designed to enhance User Experience (UX) testing by utilizing visual understanding for intelligent navigation. This approach allows for more comprehensive testing of user workflows while adapting to interface changes, addressing the limitations of traditional automated testing methods.
Flo Health Implements AI Content Review System with Amazon Bedrock
Flo Health's engineering team has developed a production-grade medical content review system using Amazon Bedrock, resulting in a 60% reduction in review time and tripled content throughput. This advancement addresses the challenges of scaling medical content production while maintaining rigorous accuracy standards without the need for expanding the medical team.
ScienceSoft launches HIPAA-compliant AI voice scheduler using AWS
ScienceSoft has developed a HIPAA-compliant AI voice scheduler leveraging Amazon technologies to streamline healthcare appointment scheduling. This solution addresses inefficiencies in traditional manual workflows, potentially transforming patient scheduling in healthcare settings.
Henry Schein One implements AI for real-time dental image verification
Henry Schein One launched Image Verify, an AI system using Amazon SageMaker, to assess dental X-ray quality in real-time. This innovation addresses a significant bottleneck in dental claims processing, minimizing delays and improving patient care as the system scales to 40,000 locations globally.
Amazon Quick Automate Introduces Native Case Management for Workflows
Amazon Quick Automate now includes native case management, enhancing the tracking and management of workflows across AI agents. This feature enables organizations to monitor work item progress, handle exceptions, and allow human intervention when necessary, facilitating efficient enterprise-scale automation.
Amazon SageMaker HyperPod Enhances LLM Inference with New Features
Amazon SageMaker HyperPod has introduced new features to enhance large language model inference, including Disaggregated Prefill and Decode, support for Hugging Face, NVMe integration, and model quantization with Unsloth. These updates aim to improve performance, observability, and resource efficiency for enterprise users.
BYOKG and GraphRAG initiatives enhance pharmaceutical drug discovery
Pharmaceutical researchers are adopting BYOKG and GraphRAG to improve access to fragmented knowledge systems. These solutions aim to integrate data sources, enhancing efficiency in drug discovery.
Amazon Bedrock Offers AI-Powered Email Management for Public Sector
Amazon Bedrock introduces an AI-driven email management solution for public sector organizations to automatically classify and prioritize incoming communications. This system enhances responsiveness and optimizes resource allocation in local government settings where email volume is high.
Amazon QuickSight Introduces Multi-Dataset Topics and Relationships for Enhanced Analytics
Amazon QuickSight has launched Multi-Dataset Topics and Multi-Dataset Relationships, allowing analytical questions across multiple datasets without the need for pre-joining them into a single table. This new capability enables AI-generated SQL queries and supports seven distinct data modeling patterns, facilitating more efficient data analysis and management. The updates are part of efforts to streamline dataset management and enhance data governance.
Amazon Nova introduces selective unlearning for customizable content moderation
Amazon Nova has launched Customizable Content Moderation Settings (CCMS), allowing clients to selectively adjust model safeguards. This development is significant as it enables organizations to generate needed content while adhering to essential AI safety policies.
AWS Introduces Multi-Turn Reinforcement Learning in SageMaker AI for Enterprise Agents
AWS has introduced new capabilities for multi-turn reinforcement learning in Amazon SageMaker AI and Amazon Nova. These developments facilitate the training of AI agents that can handle complex workflows and multi-step tasks, improving their reliability and performance. This advancement marks a significant step in enhancing enterprise AI solutions such as support ticket resolution and content moderation.
Amazon Nova automates PII redaction in images for compliance
Amazon has introduced Nova, a foundation model that automates the redaction of PII in images. This tool addresses compliance challenges by ensuring sensitive information is accurately identified and redacted, reducing the risk of regulatory penalties for organizations sharing images containing potential PII.
Amazon Bedrock introduces tools to combat AI-generated phishing risks
Amazon Bedrock offers capabilities to detect and address AI-generated phishing, adapting to sophisticated attacks. This response is crucial as traditional phishing filters fail against today's contextually accurate threats.
HippoRAG Framework Integrates Amazon Services for Enhanced Knowledge Retrieval
HippoRAG is a new Retrieval Augmented Generation framework inspired by human memory systems. It utilizes Amazon Bedrock, Neptune, and personalized PageRank for efficient multi-hop reasoning and knowledge integration across documents.
BoltzGen Deploys on Amazon SageMaker AI for Protein Design
BoltzGen is now available on Amazon SageMaker AI, optimizing protein binder design processes by managing GPU infrastructure automatically. This integration allows researchers and developers to focus on design rather than infrastructure, significantly reducing operational overhead and costs in protein design projects.
Amazon Bedrock enhances resilience for LLM inference in production environments
Amazon Bedrock introduces resilience patterns to improve large language model (LLM) inference. These patterns focus on maintaining availability, response time, and cost-effectiveness, critical as generative AI scales in production.
Outpost VFX Achieves 8x Faster AI Training Using AWS
Outpost VFX improved their AI model training speeds by 8x by leveraging AWS infrastructure. This enhancement optimizes face replacement workflows, reducing project delays and improving overall efficiency in visual effects production.
IBS Software Develops Bilingual NER with Amazon Bedrock for Cargo Logistics
IBS Software created a bilingual Named Entity Recognition (NER) system utilizing Amazon Bedrock to process cargo logistics emails in English and Japanese. This solution achieves over 95% accuracy while reducing operational costs significantly, addressing challenges in manual interventions and accuracy trade-offs.