Perplexity Search debuts at the top of the Artificial Analysis Search Index, with all three context size variants taking top positions on the leaderboard The Perplexity Search API comes with three context settings (low, medium, and high) that control how much extracted content each search result carries. We tested all three variants using our standardized methodology: the same model (GPT-5.6 Luna at medium reasoning), running inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web. Only the provider behind the search tool changes. Key results: ➤ Perplexity Search (medium) scores 80 on the Artificial Analysis Search Index, ahead of the previous leaders, Parallel (advanced) and Brave Search (LLM context), at 75. The high and low variants score 79 and 77 respectively. Its lead is concentrated in BrowseComp results, with AA-Omniscience and DeepSearchQA scoring comparably to other leading providers ➤ Efficient search payloads: smaller overall search results mean the model reads less per task, so Perplexity has the lowest model inference cost per task of providers we’ve tested so far, ranging from $0.028 to $0.034 across the three variants vs $0.036 for the next lowest provider ➤ Total cost per task is ~$0.091 for the medium and high context variants, at mid-pack latency. For comparison, Parallel (advanced) costs $0.084 per task and Brave (LLM context) costs $0.13 per task
Artificial Analysis
Technology, Information and Internet
San Francisco, California 35,411 followers
Independent analysis of AI: Understand the AI landscape and analyze AI technologies http://artificialanalysis.com/
About us
Leading independent analysis of AI. Backed by Nat Friedman, Daniel Gross and Andrew Ng.
- Website
-
https://artificialanalysis.ai
External link for Artificial Analysis
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Headquarters
- San Francisco, California
- Type
- Privately Held
Locations
-
Primary
Get directions
101 Montgomery St
500
San Francisco, California 94104, US
Employees at Artificial Analysis
Updates
-
Agnes AI's Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, up 9 points from Agnes 2.5 Pro Alpha, driven by large agentic gains but using ~2x the output tokens Agnes AI is a Singapore-based AI lab that trains full-modality foundation models and offers them through a free omni-modal API. Agnes 2.5 Pro Beta is a beta checkpoint of its next 2.5 Pro release, a text, image, and video input reasoning model with text output. At 49 on the Intelligence Index, Agnes 2.5 Pro Beta moves Agnes from mid-pack to the frontier-adjacent tier, just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52) and ahead of MiniMax-M3 (45). Key results: ➤ Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, a 9-point jump from Agnes 2.5 Pro Alpha (40). This places it just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52), and ahead of MiniMax-M3 (45). ➤ Agentic capabilities drive the jump. The Artificial Analysis Agentic Index rises from 25 to 44, just behind Gemini 3.7 Flash (high, 45) and ahead of Gemini 3.5 Flash (high, 40) and MiniMax-M3 (36). τ³-Banking nearly triples from 12% to 36%, and GDPval-AA v2 rises from an Elo of 1171 to 1456 against a human baseline of 1000. ➤ Frontier reasoning evaluations improves more modestly. Humanity's Last Exam rises from 34% to 38%, GPQA Diamond from 88% to 91%, and CritPt from 11% to 16%. ➤ The AA-Omniscience improvement from -25 to -11 comes from abstention, not increased accuracy. Agnes 2.5 Pro Beta attempts only 45% of questions against Agnes 2.5 Pro Alpha's 94%, cutting the hallucination rate from 88% to 33%, but AA-Omniscience Accuracy halves from 33% to 17%. ➤ The intelligence gain required ~2x as many tokens from its predecessor. Agnes 2.5 Pro Beta uses 50k output tokens per Intelligence Index task, more than double Agnes 2.5 Pro Alpha's 24k. Additional model details: ➤ Context window: 1M tokens ➤ Max output tokens: 65k ➤ Input modalities: Text and image ➤ Pricing: $0.10 / $0.30 / $0.01 per 1M input / output / cache hit tokens ➤ Availability: Agnes AI first-party API See the full results at https://lnkd.in/gJ5tG4Pg
-
-
MiniMax H3 Max, a post-trained version of MiniMax H3 developed by fal, debuts at #1 in Image to Video and #3 in Text to Video on the Artificial Analysis Video Leaderboards with Audio, ahead of the base MiniMax H3 on both MiniMax H3 Max is built and served by fal, and post-trained from MiniMax H3. fal describes it as being tuned for stronger prompt adherence and better aesthetics, co-optimized with their custom inference stack for higher throughput. It generates 5 to 15 second clips with native audio at up to 768p. In the Artificial Analysis Video Arena, H3 Max ranks #1 in Image to Video with Audio, narrowly ahead of ByteDance's Dreamina Seedance 2.0 720p. It ranks #3 in Text to Video with Audio, on a board where the top three models sit within 6 points of each other. fal prices MiniMax H3 Max at $0.04 per second of 768p video ($2.40 per minute). The base MiniMax H3 endpoint on fal is $0.06 per second at the same resolution. fal has stated its intent to release the weights for MiniMax H3 Max. If it does, H3 Max would become the highest ranked open weights model on both boards, ahead of MiniMax H3, which leads on open weights today.
-
-
Google has released Gemini 3.5 Transcribe, ranking #5 on AA-WER at 2.6%, alongside Gemini 3.5 Transcribe Live, achieving 4.0% AA-WER Streaming at 0.40s after speech end Google DeepMind's Gemini 3.5 Transcribe release comprises two API offerings: Gemini 3.5 Transcribe Live for continuous streaming through the Live API, and Gemini 3.5 Transcribe for pre-recorded audio through the Interactions API. Both support 85+ languages, custom vocabulary and automatic text formatting, while the pre-recorded API adds timestamped multi-speaker identification for up to three speakers, with support beyond three currently experimental. Key takeaways ➤ Non-streaming transcription: Gemini 3.5 Transcribe achieves 2.6% AA-WER, ranking #5 overall, and processes audio at approximately 84× realtime. ➤ First Final Transcription: Gemini 3.5 Transcribe Live achieves 4.0% WER, with its first final-denoted transcript arriving 0.40s after VAD-detected end of speech. ➤ First Partial Transcription: Gemini 3.5 Transcribe Live achieves 5.8% WER, with its first transcript-bearing event arriving 0.25s after detected end of speech. ➤ Price: Gemini 3.5 Transcribe costs approximately $5 per 1,000 minutes and Transcribe Live $9, assuming 25 audio tokens per second and 175 text tokens per minute at their respective API rates ($2 per 1M audio input tokens and $12 per 1M text output tokens for Transcribe; $3.50 and $21, respectively, for Live).
-
-
GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index. At $0.09 Cost per Task, it sits comfortably on the Intelligence vs. Cost per Task Pareto frontier Z AI has released GLM-5.3-Flash, a smaller and cheaper sibling to GLM-5.3 at 320B total parameters and just 18B active parameters. GLM-5.3-Flash supports low/high/max reasoning efforts, and scores 57 evaluation on the Artificial Analysis Intelligence Index with max reasoning effort. This places the model only 3 points behind GLM-5.3 at 60 and in line with GPT-5.6 Terra and Muse Spark 1.2. On Z AI's first-party API, GLM-5.3-Flash is priced at $0.15 / 1M input tokens and $0.50 / 1M output tokens, just over 10% of the price of GLM-5.3. Cached input tokens are priced at $0.026 / 1M tokens, an 80% discount. Its Cost per Task on the Intelligence Index is $0.09, compared to $0.68 for GLM-5.3 (max), and it sits on the Pareto frontier for Intelligence vs. Cost per Task. Key results: ➤ GLM-5.3-Flash is 3 points behind GLM-5.3 (max) on the Artificial Analysis Intelligence Index, at ~7.5x lower Cost per Task. At $0.09 per Intelligence Index task against $0.68 for GLM-5.3, it sits on the Pareto frontier for Intelligence vs. Cost per Task. It ties GPT-5.6 Terra ($0.51) and Muse Spark 1.2 ($0.40) at 57 while costing ~5.7x and ~4.4x less per task. ➤ GLM-5.3-Flash is less token efficient, but its low per-token pricing means this does not translate into a high Cost per Task. The model used 149M output tokens to run the Intelligence Index, ~11% fewer than GLM-5.3 at 168M, but more than Kimi K3 (133M) and Qwen3.8 2.4T A95B (136M) which score the same on the Intelligence Index. Reasoning tokens account for 134M of the 149M total (~90%). ➤ GLM-5.3-Flash matches GLM-5.3 on real-world agentic work on GDPval-AA v2. With an Elo of 1770, the model is tied within the margin of error for GLM-5.3 and Grok 4.6. This places it behind only Claude Opus 5 (xhigh and max). On Terminal-Bench v2.1 it also matches GLM-5.3 (84.3% vs 83.9%), and on τ³-Banking it trails by 3.1 p.p. at 47.2%. ➤ GLM-5.3-Flash demonstrates good real-world knowledge and hallucination rate, scoring +7 on AA-Omniscience. Its AA-Omniscience Accuracy is 28%, 6 p.p. below GLM-5.3 (max) at 34% and well below GPT-5.6 Terra at 47%. However, with a Hallucination Rate of 28%, it is an improvement over GLM-5.3 at 30%. In real-world knowledge, GLM-5.3-Flash knows less than the bigger models and frontier proprietary models in its Intelligence Index tier with an accuracy of 28%. Additional model details: ➤ Pricing: On Z AI's first-party API, $0.15 / 1M input tokens and $0.50 / 1M output tokens . Cached input tokens are priced at $0.03 / 1M tokens, an 80% discount. ➤ Accessibility: Accessible through Z AI's first-party API at launch. ➤ Size: 320B total parameters with 18B active parameters ➤ License: MIT ➤ Context Window: 1M
-
-
Breeze TTS 2 is now the leading Open Weights TTS model in the Artificial Analysis Provider Voices Speech Arena, surpassing Fish Audio S2 Pro by 90 Elo points Breeze TTS 2 is the latest TTS model from BreezeBlue, supporting 50 languages, voice generation from text prompts, and streaming generation. Its weights are openly available on Hugging Face. Key takeaways: ➤ Provider Voices: Breeze TTS 2 ranks #1 among Open Weights TTS models, leading the next best Open Weights model, Fish Audio S2 Pro at 1,125, by 90 Elo points. It also ranks #6 overall out of 100+ models, with an Elo of 1,215. ➤ Controlled Voices: Breeze TTS 2 ranks #3 among Open Weights TTS models, with the same Elo as Fish Audio S2 Pro at 1,002. It trails the leading Open Weights model, Mistral's Voxtral TTS, at 1,010 by 8 Elo points, and ranks #16 out of 39 models overall with an Elo of 1,002. ➤ Speed: Breeze TTS 2 processes 45 characters per second, trailing the leading Open Weights model, Fish Audio S2 Pro, at 102 characters per second. ➤ Price: Breeze TTS 2 is priced at $34 per 1M characters on BreezeBlue's hosted endpoint, more expensive than competitor Open Weights model, Fish Audio S2 Pro, at $15 per 1M characters, though weights are also available for self-hosting both models.
-
-
South Korea has once again established a clear #3 position in the global AI race, trailing only the United States and China. The depth and vibrancy of Korea’s AI ecosystem is supported by strong domestic talent, public investment, and the incentives created by Korea’s Sovereign AI Foundation Model project A number of South Korean labs, including Motif Technologies, Upstage, SK Telecom, and LG AI Research, have now scored above 30 on the Artificial Analysis Intelligence Index. This concentration of competitive model developers highlights the growing density of Korea’s domestic AI ecosystem. Among them, Motif Technologies and Upstage have models scoring above 40, making them the highest-scoring models on the Intelligence Index developed outside the United States and China. Leading models from each lab include: ➤ Motif Technologies’ Motif 3 scored 47 on the Intelligence Index (314B total parameters, 13B active; open weights). ➤ Upstage’s Solar Pro 4 scored 42 on the Intelligence Index. It is a proprietary reasoning model. ➤ SK Telecom’s A.X K2 scored 35 on the Intelligence Index (692B total parameters, 33B active; open weights). ➤ LG AI Research’s K-EXAONE 2.0 scored 31 on the Intelligence Index (750B total parameters, 37B active; open weights). Outside the competition, Korea’s model-development ecosystem has continued to broaden. Trillion Labs, Korea Telecom (KT), and other AI labs are building out their own model families, adding to a domestic field that is considerably deeper and more competitive than it was a year ago.
-
-
Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text to Image MAI-Image-2.6 is the newest model in Microsoft AI's MAI-Image family, announced August 10. Microsoft highlights stronger text rendering, better portraits and 3D imagery, and more polished commercial and photorealistic outputs. Like the MAI-Image-2.5 family, it handles both text to image generation and image editing. In the Artificial Analysis Image Arena, MAI-Image-2.6-Preview debuts at #1 on the Image Editing Leaderboard, ahead of Microsoft's own MAI-Image-2.5-Pro, Reve 2.1, and OpenAI's GPT Image 2, and giving Microsoft the top two spots on the board. In Text to Image it takes #2, behind only OpenAI's GPT Image 2 and ahead of Reve 2.1. On our refreshed Text to Image taxonomy, MAI-Image-2.6-Preview takes the top spot on 5 of the 19 category leaderboards (Material, Knowledge, Frontier, Retail & Ecommerce, and Marketing & Advertising), ahead of GPT Image 2, which leads every other category. MAI-Image-2.6 extends a rapid run of strong image releases from Microsoft AI. MAI-Image-2.5 debuted at #2 in Text to Image in June. MAI-Image-2.5-Pro, launched July 23, took #1 in Image Editing when we published our results last week. MAI-Image-2.6 now takes that top spot from its own sibling, and sits at #2 in Text to Image against MAI-Image-2.5-Pro's #8. MAI-Image-2.6 is available in the MAI Playground and in Private Preview on Microsoft Foundry.
-
-
Which model should you run on the iPhone 17 Pro? Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with Liquid AI As a start, we are covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use Initial results: ➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models ➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time ➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights Read more in our launch article: https://lnkd.in/gitBqcSj
-
-
You can now benchmark Qwen3.8 27B with Optima. See how a model that can run on your laptop compares on performance, cost and speed for your custom use cases We're also offering credits and free advisory from an Artificial Analysis team member to enterprises looking to build benchmarks for their use cases. Let us know if you’re interested: https://lnkd.in/gGv7S5sX Build and run your custom benchmarks today at https://lnkd.in/e_5UQhCb