1 पॉइंट द्वारा GN⁺ 2 시간 전 | 1 टिप्पणियां | WhatsApp पर शेयर करें
  • मूल्यांकन किए गए 170 मॉडलों में Claude Opus 5 Adaptive Reasoning·Max Effort ने Intelligence Index में 61 अंक के साथ पहला स्थान हासिल किया, जबकि Xhigh Effort और Claude Fable 5 ने क्रमशः 60 अंक के साथ उसके बाद स्थान पाया
  • Intelligence Index v4.1 एजेंट कार्य, कोडिंग, वैज्ञानिक reasoning, knowledge reliability और long-context reasoning को कवर करने वाले 9 evaluations को जोड़ता है, और reasoning models के extended thinking time को भी performance measurement में शामिल करता है
  • output speed में Mercury 2 सबसे तेज़ 901.6 tokens/s रहा, first-token latency में Gemini 2.5 Flash-Lite सबसे कम 0.34 सेकंड रही, और Llama 4 Scout अधिकतम 10M tokens की context window देता है
  • pricing की तुलना सिर्फ token unit price से नहीं, बल्कि input, cache hit, cache write, reasoning और answer tokens को वज़न देकर निकाले गए cost per task से की जाती है; cache write और storage cost provider के अनुसार अलग-अलग हैं
  • open-weight models में GLM-5.2 (max) 51 अंकों के साथ सबसे ऊपर है, लेकिन overall नंबर 1 से 10 अंक पीछे है; इसलिए model selection में intelligence के साथ cost, speed, latency और context window भी देखनी चाहिए

Claude Opus 5 की intelligence ranking

  • Claude Opus 5 (Adaptive Reasoning, Max Effort) को Intelligence Index में 61 अंक मिले, और यह मूल्यांकन किए गए 170 मॉडलों में नंबर 1 है
  • top 5 models इस प्रकार हैं
    • Claude Opus 5 (Adaptive Reasoning, Max Effort): 61 अंक
    • Claude Opus 5 (Adaptive Reasoning, Xhigh Effort): 60 अंक
    • Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback): 60 अंक
    • GPT-5.6 Sol (max): 59 अंक
    • Claude Opus 5 (Adaptive Reasoning, High Effort): 59 अंक
  • Max Effort configuration, 126 reasoning models में भी नंबर 1 है; reasoning models जवाब से पहले जटिल समस्याओं को संभालने के लिए extended thinking करते हैं

Intelligence Index v4.1 की evaluation संरचना

  • Intelligence Index v4.1 निम्न 9 evaluations को जोड़ता है
    • GDPval-AA v2: एजेंट-आधारित वास्तविक कार्य
    • 𝜏³-Banking: एजेंट tool use
    • Terminal-Bench v2.1: एजेंट coding और terminal use
    • SciCode: coding
    • Humanity's Last Exam: reasoning और knowledge
    • GPQA Diamond: scientific reasoning
    • CritPt: physics reasoning
    • AA-Omniscience: knowledge accuracy और non-hallucination rate
    • AA-LCR: long-context reasoning
  • उपयोग के अनुसार, overall intelligence score से ज़्यादा specific evaluation results महत्वपूर्ण हो सकते हैं
  • अलग evaluation items में निम्न benchmarks भी शामिल हैं
    • AA-Briefcase: एजेंट knowledge work
    • AutomationBench-AA: एजेंट SaaS workflow
    • Harvey LAB-AA: legal agent tasks
    • EnterpriseOps-Gym-AA: एजेंट business operations
    • IFBench: instruction following
    • APEX-Agents-AA: long-horizon agent tasks
    • ITBench-AA: Kubernetes incident root cause analysis
    • MMMU-Pro: visual reasoning

knowledge work, reliability और openness metrics

  • AA-Briefcase Elo analysis quality Elo, presentation Elo और evaluation-criteria pass rate को जोड़ता है
    • pass rate को synthetic head-to-head मुकाबलों के माध्यम से Elo में बदला जाता है
    • Elo और 95% confidence interval की boundary को 0 से नीचे नहीं जाने दिया जाता
  • AA-Omniscience Index knowledge reliability और hallucination को -100 से 100 अंकों पर मापता है
    • सही जवाबों पर इनाम मिलता है और hallucination पर penalty लगती है, लेकिन जवाब देने से मना करने पर penalty नहीं है
    • 0 अंक का मतलब है सही और गलत जवाबों की संख्या बराबर है, और negative score का मतलब है गलत जवाब सही जवाबों से ज़्यादा हैं
  • Openness Index model की openness को 0~100 normalized score पर मापता है; score जितना अधिक, model उतना अधिक open है
  • weight disclosure को अलग से दिखाया जाता है, और जिन मॉडलों के weights public हैं लेकिन commercial use के लिए paid license आदि चाहिए, उन्हें Commercial Use Restricted के रूप में चिह्नित किया जाता है

open-weight models ranking

  • मूल्यांकन किए गए 170 मॉडलों में 94 open-weight models हैं
  • open-weight models में Intelligence Index के top scores इस प्रकार हैं
    • GLM-5.2 (max): 51 अंक
    • MiniMax-M3: 44 अंक
    • DeepSeek V4 Pro (Reasoning, Max Effort): 44 अंक
  • GLM-5.2 (max) ने open-weight models में सबसे अधिक score दर्ज किया

output speed और task processing time

  • Mercury 2 की output speed 901.6 tokens/s के साथ सबसे तेज़ है
    • Gemini 3.5 Flash-Lite: 435.1 tokens/s
    • HyperNova 60B 2605: 427.6 tokens/s
    • Granite 4.0 H Small भी top tier में शामिल है
  • output speed, streaming models में पहला chunk मिलने के बाद model द्वारा token generation के दौरान प्राप्त tokens per second है
  • Intelligence Index में per-task time, प्रति task output tokens को output speed से विभाजित करके और हर benchmark का relative weight लागू करके निकाला जाता है
    • इसमें time to first token (TTFT) और अन्य overhead शामिल नहीं हैं
  • यदि अपना API मौजूद है तो उसी API का performance लिया जाता है, और Meta Llama जैसे models जिनका अपना API नहीं है, उनके लिए कई providers का median उपयोग होता है

latency और end-to-end response time

  • पहला जवाब token सबसे जल्दी देने वाला model Gemini 2.5 Flash-Lite (Non-reasoning) है, जिसकी latency 0.34 सेकंड है
    • Command A+: 0.42 सेकंड
    • North Mini Code: 0.49 सेकंड
    • Gemini 2.5 Flash भी कम latency वाले मॉडलों में शामिल है
  • first-token latency, API request के बाद पहला जवाब token मिलने तक का समय है
    • reasoning models में जवाब से पहले का thinking time भी शामिल है
    • जो models streaming support नहीं करते, उनमें पूरा response मिलने का समय लागू किया जाता है
  • 500-token end-to-end response time में निम्न तत्व जोड़े जाते हैं
    • पहला response token मिलने तक का input time
    • reasoning model द्वारा जवाब से पहले उपयोग किया गया thinking time
    • output speed के आधार पर 500 answer tokens generate करने का समय
  • thinking tokens की संख्या, अलग-अलग 60 prompts पर मापे गए औसत से ली जाती है

pricing और cost per task

  • pricing comparison में cache hit, input और output tokens के प्रति 1 million tokens डॉलर रेट का उपयोग होता है
  • सूची में Devstral 2 और North Mini Code को प्रति 1 million tokens $0.00 के रूप में दिखाया गया है, जिनके बाद Gemma 3 4B और Gemma 3 27B आते हैं
  • mixed rate के आधार पर सबसे सस्ते models निम्न हैं
    • Nova Micro: $0.03/1M tokens
    • Sarvam 30B (high): $0.03/1M tokens
    • Gemma 4 E4B (Non-reasoning): $0.03/1M tokens
  • Intelligence Index का cost per task, हर evaluation के input, cache hit, cache write, reasoning और answer token prices को tasks की संख्या से विभाजित करके और Index के evaluation weights लागू करके निकाला जाता है
  • पूरे Index को चलाने की लागत, repeat runs को छोड़कर हर evaluation की token usage और हर token type की pricing से गणना की जाती है
  • दिखाई गई cache mixed price में केवल cache hit cost शामिल है; बाकी cache costs provider के अनुसार अलग हैं
    • Anthropic cache write cost अलग से लेता है, और 5 मिनट तथा 1 घंटे TTL की rates अलग हैं; 1 घंटा अधिक महंगा है
    • Google Vertex/Gemini, cache hit pricing के अलावा प्रति घंटा storage cost भी लेता है
    • कुछ providers 200K tokens से बड़े prompts पर tiered pricing लागू करते हैं
    • OpenAI और DeepSeek आदि आम तौर पर write/storage cost के बिना केवल cache hit pricing लेते हैं

token usage और context window

  • प्रति task output tokens की संख्या, हर evaluation के output tokens को Intelligence Index के relative weights से गुणा करके और repeats को छोड़कर tasks की संख्या से विभाजित करके निकाली जाती है
  • context window के top models इस प्रकार हैं
    • Llama 4 Scout: 10M tokens
    • Grok 4.20 0309: 2M tokens
    • Gemini 1.5 Pro (May)
    • Grok 4.1 Fast
  • context window input और output tokens को मिलाकर अधिकतम सीमा है; output token limit आम तौर पर इससे बहुत छोटी होती है और model के अनुसार बदलती है
  • बड़ी context window, बड़े data से जानकारी खोजने और reasoning करने वाले RAG workflows से जुड़ी है

open-weight models का आकार

  • model size को training योग्य weights और bias के कुल total parameters तथा वास्तविक inference में चलने वाले active parameters में विभाजित किया जाता है
  • Mixture of Experts models में routing mechanism हर token पर केवल कुछ experts चुनता है, इसलिए active parameters कुल से कम होते हैं
  • Dense models सभी parameters का उपयोग करते हैं, इसलिए active parameters और total parameters समान होते हैं

comparison scope और उपयोग का तरीका

  • models की तुलना intelligence, pricing, output speed, first-token latency, end-to-end response time और context window size जैसे कई dimensions में की जाती है
  • performance metrics को 586 models पर standardized prompts के साथ सीधे मापा गया है
  • हर model के detail page पर विस्तृत metrics और समान मॉडलों के साथ direct comparison देखा जा सकता है, और model selector से chart में दिखने वाले models को समायोजित किया जा सकता है

1 टिप्पणियां

 
GN⁺ 2 시간 전
Hacker News की रायें
  • यह leaderboard अब end users के लिए यह तय करने में लगभग बेकार हो गया है कि कौन-सा model कब इस्तेमाल करना है
    हर model की अलग-अलग domains और tasks में strengths और weaknesses अलग होती हैं, इसलिए एकल metric ranking उपयोगी नहीं है; अब कोई एक ‘सबसे अच्छा model’ भी नहीं है, और task की complexity के हिसाब से शायद सबसे अच्छे model की ज़रूरत भी न पड़े
    उदाहरण के लिए UI design के लिए Fable, backend system design के लिए Sol, vulnerability development के लिए Kimi K3 इस्तेमाल किया जा सकता है; ऐसे metrics आखिरकार model companies के दिखावे वाले numbers के ज्यादा करीब हैं

    • एक साल पहले नया ‘सबसे अच्छा model’ आया था, तो उम्मीद के साथ उसे 3 files में मामूली बदलाव वाला एक simple programming task दिया, और उसने ठीक कर दिया
      लेकिन उसी family का पुराना और छोटा model भी उसे ठीक से कर गया, वह 3 गुना तेज़ और लागत में 1/9 था; उसी क्षण बात समझ आ गई
    • tech stack-wise benchmarks हों तो अच्छा होगा
      जैसे 2026 में Elixir/Phoenix project को सबसे idiomatic तरीके से कौन-सा model संभालता है—भले ही यह थोड़ा subjective हो, फिर भी जब सभी models की लगातार खुद तुलना करना मुश्किल है और performance difference भी बड़ा है, तो यह उपयोगी हो सकता है
    • ‘पूरी तरह बेकार’ या ‘एकमात्र उद्देश्य’ कहना अतिशयोक्ति है
      हर statistic में distortion होता है, लेकिन कुछ भी न जानने से बेहतर है; benchmark कई tasks का average दिखाता है, यही statistics का मूल उद्देश्य—summary—है
    • link खोलकर देखें तो सिर्फ single metric नहीं है, बल्कि decision के लिए ज़रूरी सभी metrics वास्तव में दिए गए हैं
    • Artificial Analysis के cost per task या hallucination frequency जैसे charts में value है
      लेकिन इन्हें एक result में मिला देने पर सारी nuance गायब हो जाती है और सिर्फ दिखावे वाली ranking बचती है
  • अगर मैं Google executive होता, तो सिर्फ काम ज़्यादा होने की वजह से नहीं, बल्कि लोगों के सामने दिखने में शर्म आने की वजह से भी office में ही सोना पड़ता, ऐसा लगता है

    • subsidies जलाते हुए घाटा न करने वाला इकलौता player है, इसलिए उल्टा चैन से सो सकता है
    • Google शायद सबसे intelligent AI से ज़्यादा सभी products में monetize कर पाने की speed को महत्व देता है
      वह बहुत पहले से कहता आया है कि सवालों के सही जवाब सबसे तेज़ देना उसकी top priority है
    • यह भी संदेह है कि Google OpenAI या Anthropic की तरह intelligence-maximized models से compete करना चाहता भी है या नहीं
      तीनों में वह AGI की पूजा करते हुए तपस्या न करने वाली इकलौती company जैसा दिखता है
    • GPQA Diamond benchmark में joint #1 है: https://artificialanalysis.ai/evaluations/gpqa-diamond
      शायद उसका लक्ष्य सिर्फ यही हो
    • Gemini models भी कई categories में lead कर रहे हैं या top tier में हैं, इसलिए यह conclusion सही नहीं लगता कि वे शर्मनाक रूप से पीछे हैं
  • intelligence index के top में Claude Opus 5 Max 61 points, Opus 5 Xhigh 60 points, Claude Fable 5 Max और Opus 4.8 replacement model combination 60 points, GPT-5.6 Sol Max 59 points, Opus 5 High 59 points हैं
    इसलिए Opus 5 Xhigh, Sol Max से ऊपर है, और Opus 5 High, Sol Max के बराबर है; ऐसे में दोनों models की price और speed difference जानने की उत्सुकता है

    • Artificial Analysis के intelligence बनाम cost और time per task graphs में Opus 5 High और Sol Max की cost और time लगभग समान हैं
      DeepSwe में Opus 5, Fable को हराता है लेकिन Sol से पीछे रहता है; FrontierCode में सबको dominate करता है, पर reasoning effort को Medium से ऊपर करने पर Sonnet level तक तेज़ी से गिर जाता है
    • Opus 5, Sol Max जितना intelligent नहीं है और कुछ tasks में Opus 4.8 से भी खराब दिखता है
      खासकर यह superficial तरीके से approach करता है, app या framework को पूरा सिखाना पड़ता है, और पहले से किसी दूसरे रूप में supported feature को फिर से जोड़ने जैसे मूर्खतापूर्ण काम से शुरू करता है
  • components में AA-Omniscience Index दिलचस्प है
    यह knowledge reliability और hallucination को मापता है, सही जवाबों को reward करता है और hallucination पर penalty देता है, लेकिन answer refusal पर penalty नहीं देता; यह parameter scale या density को मोटे तौर पर दिखाने वाला metric लगता है
    ranking है: Claude Fable 5 replacement model combination, Gemini 3.1 Pro Preview, Claude Opus 5 Max, Grok 4.6 High, Gemini 3.6 Flash, GPT-5.6 Sol Max; और Gemini 3.x में काफी समय से large model वाली खास feeling आती रही है

    • Gemini 3.1 Pro knowledge tasks में बहुत अच्छा है और Google ने इस हिस्से को अच्छे से किया है
      Gemini 3.6 Flash की image analysis भी solid है, लेकिन coding या agent tasks में कमी है
    • 3.6 Flash का 24 points के साथ Sol के 22 से ऊपर होना देखकर यह पक्का कहना मुश्किल है कि यह parameter scale का proxy metric है
      3.x Flash के बारे में कहा जाता है कि यह single TPU 8i में फिट हो जाता है, और इसकी processing speed भी 234.7 tokens/sec है, जो Sol के 64.4 और Opus के 56.3 से भारी बढ़त में है; 3.1 Pro भी 113.9 tokens/sec के साथ कहीं तेज़ है
      AA-Omniscience Accuracy में super-large Fable 61% के साथ आगे है, और Sol, GPT-5.5, Opus जैसे बड़े frontier models उसके पीछे हैं, जो expectation से ज्यादा मेल खाता है
      लगता है Gemini ने coding top score की बजाय general knowledge और operating cost efficiency पर focus किया है; normalized category scores में सिर्फ software में competitiveness कम है
    • अगर answer refusal पर penalty नहीं है, तो यह बहुत उपयोगी नहीं है
    • model size से ज्यादा या कम से कम उतना ही असर grounding का हो सकता है, और Google साफ वजहों से इसमें बहुत अच्छा है
  • उम्मीद करने से पहले intelligence बनाम cost matrix देखनी चाहिए: https://artificialanalysis.ai/models?intelligence-index-toke...

    • result quality तक को ध्यान में रखें तो यह हैरानी की बात है कि GPT-5.6 Sol Max सब से सस्ता है
    • Max में extra reasoning amount बहुत ज़्यादा है, इसलिए High में कितने solved tasks कम हो जाते हैं, यह जानने की उत्सुकता है
      cost काफी कम होनी चाहिए
  • Meta Muse Spark को नए नज़रिए से देखने लगा हूं
    यह किसी एक चीज़ में सबसे अच्छा नहीं है, लेकिन leaderboard में कई जगह sweet spot पर बैठता है और cost-performance balance अच्छा है
    list में न मौजूद Poolside Laguna S कहां फिट होता है, यह भी जानना चाहूंगा; निजी तौर पर मेरी दिलचस्पी ऐसे models में ज्यादा है जो अच्छे से काम करें और cost-efficient हों

  • कांटे की टक्कर में पहला स्थान पाने के बावजूद, अगर ‘safety guardrails’ के नाम पर सेंसरशिप जवाब देने से इनकार कर दे या किसी कमजोर मॉडल पर डाउनग्रेड न हो, इसका ध्यान रखना पड़े, तो उपयोगिता बहुत घट जाती है
    इसी वजह से मौजूदा workflow के कुछ हिस्सों को छोड़कर Claude का इस्तेमाल लगभग बंद कर दिया है, और 61 व 57 अंकों के फर्क से ज़्यादा विश्वसनीयता अहम है
    सेंसरशिप और वह identity verification, जिसका मैंने खुद सामना तो नहीं किया, दोनों को ध्यान में रखें तो Claude सबसे ज़्यादा प्रभावित और अस्थिर मॉडल है, और नए version का मामूली benchmark optimization इसके लायक नहीं है

    • जानना चाहूंगा कि आप किस तरह के सवाल पूछते हैं कि इतनी बार सेंसरशिप में फंस जाते हैं
    • खुद इस्तेमाल करके देखा तो यह बिल्कुल भी सिर्फ benchmarks को target करने वाला मॉडल नहीं है
      मैं सभी models पर game-making test लगाता हूं, और Opus 5 पीढ़ी बदलने जितनी बड़ी छलांग जैसा लगा; benchmarks सही तस्वीर नहीं दिखाते, इसलिए खुद test करना चाहिए
    • online limits या model में सुधार की बातें देखकर जब भी Anthropic को कभी पैसा न देने के अपने सिद्धांत पर फिर से विचार करता हूं, model card देखता हूं और उलटे मेरा विश्वास और मजबूत हो जाता है
      समझ नहीं आता Anthropic automatic, silent downgrade पर इतना अड़ा क्यों है, और संदेह है कि कोई एक भी user जवाब से इनकार के बजाय performance को अपने-आप घटाया जाना चाहता होगा
    • Claude Opus 5.0 से test environment के PaaS में global API key से web app deploy करके data लाने को कहा, तो उसने permissions बहुत broad होने की security concern बताकर मना कर दिया
      यही काम Opus 4.5~4.8 और Fable, Codex 5.4·5.5, Sol 5.6 ने बिना समस्या किया
      Opus 5 बहुत ज़्यादा सावधान है, इसलिए productive तरीके से इस्तेमाल करना मुश्किल है और इसमें और tuning चाहिए
    • शायद इसलिए कि meaningful काम नहीं किया जा रहा हो; Terence Tao ‘woke AI model’ की वजह से शिकायत नहीं करते
  • अच्छा है कि Opus 5, 4.8 की तरह हर चीज़ को फिर से explain नहीं करता
    GPT-5.6 Sol छोटी बातों पर भी बहुत ज़्यादा गहराई से reasoning करता है, जबकि Opus 5 शानदार model है और Anthropic ने अच्छा काम किया है

    • Sol को implementation spec देने पर उसने बहुत अच्छा handle किया, इसलिए यह फर्क पता नहीं चला था
      बिना spec के खुद से करने को कहा तो 2 घंटे 30 मिनट में weekly usage का 30% खर्च कर दिया, लेकिन spec देने पर 30 मिनट में सिर्फ 2% use हुआ और result भी structure, length, validation, scope और cost—सब में बेहतर था
      Sol को अपने हाल पर छोड़ दें तो वह control से बाहर हो जाता है, इसलिए अब Sol spec लिखता है और Terra implement करता है; यह तरीका अच्छा काम कर रहा है और कुल usage भी घटा है
  • ज्यादा दिलचस्प नतीजा यह है कि Opus 5, Fable 5 के बाद काफी भारी अंतर से महंगा model है
    GPT-5.6 और Kimi K3 आधी cost में 1~2% के अंतर वाले मिलते-जुलते scores देते हैं

    • chart Max reasoning effort पर आधारित है, जिसे मुख्यतः price-insensitive enterprise users इस्तेमाल करते हैं
      Medium में cost K3 के लगभग आधे तक आ जाती है, और coding tasks के 95% के लिए पर्याप्त होने की संभावना बहुत अधिक है
  • मैं किसी पक्ष में नहीं हूं, लेकिन इस result में GPT-5.6 Sol Max बेहतर दिखता है
    लगभग वही performance करीब आधी cost में देता है, और Opus 5 भी GPT-5.6 से इतना महंगा है, तो इससे साफ दिखता है कि Fable की pricing कितनी ऊंची है

    • सही तुलना करनी हो तो Opus High और Sol Max की तुलना करनी चाहिए