Amazon Science’s cover photo
Amazon Science

Amazon Science

Research Services

Seattle, Washington 393,846 followers

The latest news and research from Amazon’s science community. #AmazonScience

About us

Amazon Science gives you insight into the company’s approach to customer-obsessed scientific innovation. Amazon fundamentally believes that scientific innovation is essential to being the most customer-centric company in the world. It’s the company’s ability to have an impact at scale that allows us to attract some of the brightest minds in artificial intelligence and related fields. Our scientists continue to publish, teach, and engage with the academic community, in addition to utilizing our working backwards method to enrich the way we live and work. Follow us on LinkedIn and visit our website to get a deep dive on innovation at Amazon, and explore the many ways you can engage with our scientific community. #AmazonScience

Industry
Research Services
Company size
10,001+ employees
Headquarters
Seattle, Washington
Founded
2020
Specialties
Artificial Intelligence, Machine Learning, Computer Vision, Cloud, Economics, Sustainability, AI, ML, Conversational AI, Natural Language Processing, NLP, Robotics, Security, Privacy, Information, Knowledge Management, Operations, Scientific Research, Search, Amazon, and Alexa

Updates

  • Ten years ago, Amazon's Automated Reasoning Group set out to use mathematical logic to prove AWS systems work correctly. Today, their production services process billions of queries daily, powering tools millions of customers rely on, from IAM Access Analyzer to Amazon Inspector to Bedrock Guardrails. One project stands out: they proved correct and seamlessly replaced the AWS authorization engine, which handles one billion API calls per second, and verified it against quadrillions of production authorizations. Now the same formal-verification techniques that secured cloud infrastructure are being applied to AI, setting boundaries for autonomous agents and validating AI-generated content with up to 99% verification accuracy.

  • AWS Trainium Frontier challenges researchers to train language models from scratch on purpose-built AI chips. The NeurIPS 2026 competition is open for registration, limited to 100 teams. Teams optimize across the full stack, including model architecture, optimizer, training loop, and custom hardware kernels. No prior Trainium experience is required. Prizes include $25K for first place, co-publication with Annapurna Labs researchers, and a chance to present at an event during NeurIPS 2026 in Sydney. Register by September 30: https://amzn.to/3TL03EA

  • Announcing the 34 recipients of the Amazon Research Awards Build on Trainium program, a $110 million credit initiative supporting AI research at 30 universities including Stanford, UC Berkeley, UIUC, UCLA, CMU, and MIT. This cycle focused on Responsible AI, inviting proposals in AI safety and alignment, multi-lingual language models, representation engineering, sustainability and small language models, and deep learning models for synthetic data generation. Awardees have access to more than 700 Amazon public datasets, AI/ML services and tools, and AWS Trainium resources including tutorials and hands-on sessions.

  • AI's full potential is bottlenecked by an efficiency problem that spans the entire stack. At Berkeley RDI's Agentic AI Summit, Amazon SVP Peter DeSantis traced how the predictable memory and compute flow of AI models led Amazon to build Trainium on a systolic array architecture, stripping out flexibility those workloads don't need while preserving what matters. But no single chip will power the next decade. As AI workloads keep evolving, Peter sees a growing and more diverse hardware ecosystem ahead. Annapurna Labs' John Liu followed with a deep dive on using agentic looping to optimize models on Trainium. Key takeaways: test agents on data they've never seen, verify visibility scope before fixing multi-agent failures, and recognize when existing rules are causing the very failures you're trying to patch.

  • Multitask training objectives fight each other at every gradient step. What if you stopped forcing them to share? ControlG reframes multitask coordination as a temporal allocation problem, dedicating computational capacity to one objective at a time instead of blending conflicting gradients per step. A proportional-integral-derivative (PID) controller decides which objective needs attention next, operating across three time scales: estimating per-objective difficulty, optimizing per-epoch allocation via log-hypervolume sensitivity, and tracking the plan with feedback loops. Amazon researchers applied ControlG to the problem of graph self-supervised learning. Across nine graph datasets, ControlG achieves average ranks of 1.4, 1.9, and 1.8 for node classification, link prediction, and node clustering, exceeding all baselines.

  • Most healthcare AI benchmarks test medical knowledge in static form or evaluate agentic, tool-using agents on technical tasks done for providers. Neither captures what a patient-facing agent has to do: reason about a patient's health record over a multiturn conversation and decide what actions to take. To capture such complex interactions, PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation. An LLM-as-a-jury panel scores each conversation against over 100 clinician-vetted criteria across six dimensions — clinical safety, triage quality, workflow accuracy, task completion, clinical helpfulness, and conversational quality. The framework generates fresh scenarios on demand, preventing "training contamination", where pretrained models learn about common benchmarks from published results. Results across frontier models reveal a severity paradox: agents score higher on obvious emergencies than on routine cases that hide real risk. The dominant safety failure was crisis resource omission — recognizing suicidal ideation but failing to provide hotline information. Model capability alone narrows clinical gaps but does not close them.

    • No alternative text description for this image
  • As AI agents take on higher-stakes decisions, Amazon is further investing in mathematical proof to enable verified, trustworthy AI agents. We're providing substantial, long-term financial support to the Lean Focused Research Organization, the single largest donation in the FRO's history, to make proof accessible to every developer in the world. Mathematical proof shows with certainty that a system cannot behave incorrectly, no matter what inputs it gets. Amazon already uses Lean to help prove AI agent boundaries are correct in Bedrock AgentCore, guarantee differential-privacy protections in AWS Clean Rooms are sound, and verify compilation to Amazon's AI acceleration chips.

  • A faithful conversation transcript is not a faithful token record. Agent harnesses do useful things that make RL bookkeeping harder: compacting older messages, retrying malformed tool calls, branching into subagents, merging results back. Each rewrite is a chance for the next request's token sequence to drift from what the model generated. Turnstile is an open-source Rust proxy that records the exact token-level history at the only point where it's correct: the moment of generation. It speaks the Chat Completions API, requires no harness changes, and exports generic trajectories with token IDs, log probabilities, loss masks, and weight-version boundaries. It also captures mixture-of-experts routing decisions and processed multimodal inputs, splitting the trajectory rather than training under incorrect state.

    • No alternative text description for this image
  • The Chronos family of models has reached 1 billion downloads on Hugging Face! 🤗 Pretrained time series models have enabled inference-only forecasting systems that produce accurate predictions without task-specific training — but existing approaches largely focus on univariate forecasting, limiting their use in real-world scenarios where multivariate data and covariates matter. Chronos-2 addresses this. Our foundation model handles univariate, multivariate, and covariate-informed forecasting in a zero-shot manner, outperforming existing time series foundation models by a substantial margin across multiple benchmarks.

    • No alternative text description for this image

Affiliated pages

Similar pages