- Google ने AI एजेंट डेवलपमेंट की दक्षता, कम latency और reliability के लिए डिज़ाइन किए गए नए Gemini मॉडल 3.6 Flash, 3.5 Flash-Lite और 3.5 Flash Cyber लॉन्च किए
- Gemini 3.6 Flash, 3.5 Flash की तुलना में बेहतर coding, knowledge work और multimodal performance देता है, साथ ही output token उपयोग घटाकर लागत कम करता है
- Gemini 3.5 Flash-Lite, 3.5 class का सबसे तेज़ और सबसे cost-efficient मॉडल है, जो agentic search और document processing जैसे high-throughput कामों के लिए उपयुक्त है
- Gemini 3.5 Flash Cyber, cyber security vulnerabilities के लिए tuned एक विशेष मॉडल है, जिसे CodeMender के माध्यम से limited-access pilot program में उपलब्ध कराया जाएगा
- Gemini 3.5 Pro जल्द ही व्यापक रूप से जारी किया जाएगा और Gemini 4 का development जारी है
GN⁺ अतिरिक्त सारांश
Neo ने मूल लेख के आधार पर अतिरिक्त संक्षेप तैयार किया है
- Google ने बड़े पैमाने पर AI एजेंट बनाने के लिए ज़रूरी token efficiency, latency और stability को लक्ष्य बनाकर general-purpose, high-speed और security-specialized मॉडल अलग-अलग पेश किए
- 3.6 Flash 3.5 Flash की तुलना में 17% कम output token उपयोग करता है, reasoning steps और tool calls घटाता है, और इसकी कीमत input के 10 लाख token पर $1.50 तथा output के 10 लाख token पर $7.50 है
- 3.5 Flash-Lite प्रति सेकंड 350 output token देता है, इसकी कीमत input के 10 लाख token पर $0.3 और output के 10 लाख token पर $2.5 है, और यह thinking level को समायोजित कर bulk tasks तथा multi-step sub-agent workloads संभाल सकता है
- 3.5 Flash Cyber CodeMender की multi-agent setup में vulnerabilities को detect, verify और patch करता है, और dual-use risk को देखते हुए इसे केवल सरकारों और trusted partners तक सीमित रूप से उपलब्ध कराया जाएगा
- 3.6 Flash और 3.5 Flash-Lite, Gemini API, enterprise platform और Gemini app सहित कई जगह उपलब्ध हैं, जबकि Google 3.5 Pro partner testing और Gemini 4 pre-training भी आगे बढ़ा रहा है
तीन मॉडलों की भूमिका
- Flash series को AI एजेंट की efficiency और quality के बीच संतुलन बनाते हुए agent workflows को scale करने के लिए डिज़ाइन किया गया है
- 3.6 Flash coding, knowledge work और multimodal performance को बेहतर बनाने वाला flagship मॉडल है
- Artificial Analysis Index के अनुसार यह 3.5 Flash से 17% कम output token उपयोग करता है
- Datacurve के DeepSWE जैसे कुछ benchmarks में token उपयोग 65% तक कम हुआ
- output token की कीमत भी 3.5 Flash से कम है
- 3.5 Flash-Lite 3.5 series का सबसे तेज़ और सबसे cost-efficient मॉडल है, और agent workflows में भी पिछली Flash-Lite generation से बेहतर प्रदर्शन करता है
- 3.5 Flash Cyber cyber security-specialized model और CodeMender agent infrastructure को मिलाकर frontier models से प्रतिस्पर्धा करने लायक performance का लक्ष्य रखता है
- Gemini 3.5 Pro अभी partners के साथ testing में है और तैयार होते ही इसे व्यापक रूप से उपलब्ध कराया जाएगा
- Gemini 4 के लिए अब तक का सबसे महत्वाकांक्षी pre-training run भी शुरू हो चुका है
3.6 Flash की दक्षता और कीमत
- developers और customers से मिले 3.5 Flash feedback को ध्यान में रखकर coding, knowledge work quality और token efficiency को साथ में बेहतर किया गया है
- Artificial Analysis Index में यह 3.5 Flash की तुलना में 17% कम output token उपयोग करता है, और multi-step workflows पूरा करने के लिए reasoning steps तथा tool calls की संख्या भी घटती है
- इसकी कीमत input के 10 लाख token पर $1.50 और output के 10 लाख token पर $7.50 है
- कम token usage और कम कीमत के संयोजन से प्रति agent task कुल लागत कम होती है
- OSWorld validation API tasks में यह 3.5 Flash की तुलना में token को अधिक कुशलता से उपयोग करता है और अनावश्यक रूप से लंबे outputs घटाता है
3.6 Flash की गुणवत्ता में सुधार
- coding tasks में अनचाहे code edits और execution loops घटाकर precision बढ़ाई गई है
- DeepSWE: 49%, 3.5 Flash: 37%
- MLE Bench: 63.9%, 3.5 Flash: 49.7%
- computer use capability भी बेहतर हुई है
- OSWorld-Verified: 83.0%, 3.5 Flash: 78.4%
- computer use feature, Gemini API और Gemini Enterprise में client-side built-in tool के रूप में उपलब्ध है
- knowledge work performance भी 3.5 Flash से आगे है
- GDPval-AA v2: 1421, 3.5 Flash: 1349
- Hebbia और Harvey ने document parsing, chart और data analysis, तथा report writing जैसे multimodal tasks की क्षमता की पुष्टि की
- उपयोग के उदाहरण इस प्रकार हैं
- AIS के Managed Agents के साथ financial data और transcripts का analysis
- AGY की multi-agent orchestration के साथ latency घटाकर high-quality code migration
- Gemini app के canvas से 3D workflow के लिए photo texture extractor विकसित करना
- AGY और tldraw offline editor के साथ interactive theme studio बनाना
3.6 Flash के safety measures
- chemical, biological, radiological, nuclear (CBRN) और cyber attack misuse domains पर मजबूत Frontier Safety protections लागू किए गए हैं
- jailbreak attacks के प्रति resistance को काफी बढ़ाने के साथ उपयोगी उपयोगों पर अनावश्यक refusal कम रखने के लिए इसे train किया गया है
- अधिक जानकारी 3.6 Flash model card में उपलब्ध है
3.5 Flash-Lite की speed और cost
- इसे उन developer workflows के लिए डिज़ाइन किया गया है जिन्हें agentic search और document processing की तरह कम latency या high throughput चाहिए
- Artificial Analysis के मापदंड के अनुसार यह प्रति सेकंड 350 output token देता है, जो 3.5 series में सबसे तेज़ है
- इसकी कीमत input के 10 लाख token पर $0.3 और output के 10 लाख token पर $2.5 है, और quality 3.1 Flash-Lite से काफी बेहतर है
- यह 3.5 Flash से कम latency पर bulk tasks चला सकता है
- workload के अनुसार thinking level समायोजित किया जा सकता है
- minimal और low level, bulk tasks की latency और cost घटाने में उपयोगी हैं
- high thinking level, multi-step sub-agent workloads संभालता है
- computer use feature भी कई environments में agent tasks को support करने वाले built-in tool के रूप में उपलब्ध है
3.5 Flash-Lite के benchmarks और उपयोग
- 3.1 Flash-Lite की तुलना में coding, long context और real-world task execution performance बेहतर हुई है
- Terminal-Bench 2.1: 54%, 3.1 Flash-Lite: 31%
- GDM-MRCR v2: 72.2%, 3.1 Flash-Lite: 60.1%
- GDPval-AA v2: 1140, 3.1 Flash-Lite: 642
- कई agent और coding evaluations में यह 3 Flash से भी आगे है
- SWE-Bench Pro: 54.2%, 3 Flash: 49.6%
- OSWorld-Verified: 74.0%, 3 Flash: 65.1%
- यह 2.5 Flash और 3 Flash आधारित workloads के लिए तेज़ और अधिक सक्षम विकल्प हो सकता है
- उपयोग के उदाहरण इस प्रकार हैं
- बड़े e-commerce datasets से product attributes निकालना और summarize करना
- 3.6 Flash को top-level agent बनाकर 25 explorable web design concepts तैयार करना
- multimodal understanding के साथ receipt translation और summarization को scale करना
- कई विकल्प तेज़ी से generate और iterate करते हुए game बनाना
- अधिक जानकारी 3.5 Flash-Lite model card में उपलब्ध है
CodeMender का 3.5 Flash Cyber
- AI models मौजूदा systems की patching speed से तेज़ी से security vulnerabilities खोज सकते हैं, इसलिए software security के लिए उच्च capability और efficiency वाला approach ज़रूरी है
- Gemini 3.5 Flash Cyber को 3.5 Flash के आधार पर security vulnerability detection और remediation के लिए fine-tune किया गया है
- इसे बड़े मॉडलों की तुलना में कम token cost पर code security issues को detect, verify और patch करने के लिए डिज़ाइन किया गया है
- CodeMender में कई 3.5 Flash Cyber agents मिलकर एक unified report बनाते हैं, और CyberGym benchmark में इसने frontier-level competitive performance दर्ज की है
- तकनीक की dual-use nature को देखते हुए इसकी deployment scope सीमित रखी गई है
- CodeMender के माध्यम से इसे सरकारों और trusted partners को limited-access pilot के रूप में उपलब्ध कराया जाएगा
- इसका उद्देश्य defenders को malicious exploitation से पहले critical vulnerabilities detect और fix करने में मदद देना है, जबकि व्यापक misuse को कम रखना है
उपलब्धता
- 3.6 Flash और 3.5 Flash-Lite लॉन्च के दिन से ही कई products में उपलब्ध हैं
- developers, Gemini API को Google AI Studio और Android Studio में उपयोग कर सकते हैं
- 3.6 Flash Google Antigravity में भी उपलब्ध है
- implementation details Developer Guide में देखी जा सकती हैं
- enterprises इसे Gemini Enterprise Agent Platform में उपयोग कर सकते हैं, और 3.6 Flash Gemini Enterprise app में भी उपलब्ध है
- आम users इसे Gemini app में उपयोग कर सकते हैं, और 3.5 Flash-Lite को Google Search में भी चरणबद्ध रूप से rollout किया जा रहा है
- Google का कहना है कि user feedback को भविष्य के Gemini models में शामिल किया जाएगा और 3.5 Pro भी जल्द जारी किया जाएगा
1 टिप्पणियां
Hacker News की राय
Google छोटे मॉडल ट्रेन करने में अंदरूनी तौर पर जिन Pro मॉडल के आकार का इस्तेमाल करता है, वह कितना बड़ा है, यह जानने की जिज्ञासा है
Flash के साथ Pro जारी न होने की वजह यह हो सकती है कि मॉडल इतना बड़ा है कि वह किफायती नहीं है, या उसे सर्व करने के लिए compute resources कम हैं, या alignment से जुड़ी समस्याएँ बहुत ज़्यादा हैं
बेंचमार्क देखें तो प्रदर्शन मध्यम स्तर का है, लेकिन time-to-task intelligence और output-speed intelligence के हिसाब से यह बहुत तेज़ मॉडल है
Antigravity में इस्तेमाल किया गया 3.5 Flash भी, अगर उसका use case समझा जाए, तो कम आंका गया मॉडल है; frontend काम में यह GPT-5.5 से कहीं बेहतर था और iterative development तेज़ था, इसलिए उम्मीद है कि 3.6 भी वैसा ही होगा
Gemini बड़े context वाले पूरे MR को एक बार में संभालने में अच्छा है, लेकिन tool use और agentic coding में काफ़ी कमजोर है
लगता है Google ने AI products में सफलता सामने होते हुए भी खुद उसे गंवा दिया
बिना किसी ठीक successor product के AI Ultra subscription बंद कर दिया गया, जिससे कंपनी और मुझे Antigravity छोड़ना पड़ा। Antigravity IDE में Google Workspace advanced user subscription जोड़ा नहीं जा सकता और Gemini Enterprise Agent Platform भी integrate नहीं होता
Gemini Enterprise Agent Platform का setup process बहुत ख़राब है, और per-user spending limit लगाने के लिए हर user के लिए अलग project बनाना पड़ता है। जिन billing accounts में free credits बचे हों, उनमें Anthropic models activate न कर पाना भी अजीब है
Google और Gemini का सक्रिय समर्थक होने के बावजूद, अचानक लिए गए product फैसलों की वजह से आखिरकार Anthropic और OpenAI के 200 डॉलर प्रति माह वाले subscriptions खुद खरीदने पड़े
agy जिस endpoint को call करता है, उसे सीधे call न कर पाने से निराशा हुई, यहाँ तक कि Gemini से bypass का तरीका भी पूछा। विडंबना यह है कि Gemini, agy client के खिलाफ़ जाने में भी इतनी उत्साह से मदद करता है कि वह अब भी पसंद आता है
दूसरे models के साथ कोई तुलना ही न होना निराशाजनक है, और ऐसा भी नहीं लगता कि performance curve को ऊपर धकेला गया है। 3.6 Flash, GLM 5.2 से महँगा दिखता है जबकि प्रदर्शन कमज़ोर लगता है, लेकिन पोस्ट में details भी कम हैं
एक समय लगा था कि Google सचमुच रफ़्तार पकड़ रहा है, लेकिन समय के साथ शक बढ़ रहा है; 3.5 Pro का इंतज़ार करना होगा
3.6 Flash और 3.5 Flash-Lite के pelican results यहाँ देखे जा सकते हैं। Cyber अभी API में उपलब्ध नहीं है
LLM प्रगति में क्या बेहतर होने का मतलब अब ज़्यादा elements जोड़ना बनता जा रहा है, और क्या यह उन reasoning training का नतीजा है जो ज़्यादा tokens और ज़्यादा elements को reward करती हैं?
प्रति 10 लाख input·output tokens की कीमत इस प्रकार है
2.5 Flash $0.3/$2.5 है, 3.0 Flash $0.5/$3, 3.5 Flash $1.5/$9, और 3.6 Flash $1.5/$7.5 है
2.5 Flash-Lite $0.1/$0.4 है, 3.1 Flash-Lite $0.25/$1.5, और 3.5 Flash-Lite $0.3/$2.5 है
Gemini जिस क्षेत्र में अपनी श्रेणी से बढ़कर अच्छा करता है, वह मूलतः “Google पर यह ढूँढ दो” वाले ज्ञान-आधारित काम हैं, इसलिए इसका उपयोग-स्थान है। Google कुल मिलाकर मजबूत है, लेकिन text models में दूसरे दर्जे में है, जबकि life sciences·image·video में पहले दर्जे में है
Google ने कहा है कि वह Gemini 3.5 Pro का partners के साथ परीक्षण कर रहा है और तैयार होते ही इसे व्यापक रूप से उपलब्ध कराने की योजना है, और Gemini 4 के लिए अब तक का सबसे महत्वाकांक्षी pre-training भी शुरू कर दिया गया है
GLM-5.2 से कम बुद्धिमान, ज्यादा महँगा, और weights भी public नहीं हैं
Google models पर निर्भर रहना डरावना है। मैंने price-sensitive काम Flash 2.5 Lite पर चलाए थे, लेकिन उसका support अब खत्म हो गया है, और उसका विकल्प 3.1 Flash-Lite कहीं ज्यादा महँगा है, ऊपर से उसकी end date भी पहले से तय है
3.5 Flash-Lite इससे भी महँगा है, इसलिए कीमत लगातार बढ़ने पर भी उसके साथ चलना पड़ रहा है
प्रतिस्पर्धी कंपनियाँ अगर किसी service को बंद भी करती हैं, तो कम से कम उसे front page से हटाती हैं, पूरे docs में नोटिस लगाती हैं, और विनम्रता से कहती हैं कि नए projects में उसका उपयोग न करें
अभी भी 3.1 Pro पर अटका Jules के बारे में कोई update नहीं है। यह niche product हो सकता है, लेकिन फोन से निर्देश देकर 15 मिनट बाद GitHub PR review और merge कर पाना, commute के दौरान personal projects में काम आता था जब laptop निकालना मुश्किल हो
मैं किसी ठीक-ठाक विकल्प की तलाश में हूँ
समझ नहीं आता कि Gemini web अब तक connectors·MCP·plugins को सपोर्ट क्यों नहीं करता। गंभीर उपयोग के लिए यह शुरुआत से ही बाहर हो जाता है