2 पॉइंट द्वारा GN⁺ 2024-09-03 | 1 टिप्पणियां | WhatsApp पर शेयर करें
  • paraLLEl-GS PS2 के GS(Graphics Synthesizer) को Vulkan compute में रीक्रिएट करता है और लगभग 20 साल से de facto standard रहे GSdx की छोड़ी हुई accuracy और upscaling सीमाओं को संबोधित करता है
  • GS 4MiB VRAM और ऊंचे fillrate के आधार पर काम करता है, लेकिन destination alpha test, conditional blending, 1.0 से ऊपर alpha/color जैसी pixel pipeline विशेषताओं के कारण इसे सामान्य graphics API से मैच करना मुश्किल है
  • implementation VRAM को page और 256-byte block स्तर पर track करता है, और CLUT snapshot, texture unswizzle, render pass batching को मिलाकर framebuffer/texture feedback संभालता है
  • Tales of the Abyss, Final Fantasy X, MGS2, Valkyrie Profile 2, Shadow of the Colossus जैसे उदाहरणों में upscaling, UI, high-precision blending, texture feedback समस्याओं की तुलना की गई, और 8x/16x SSAA scenes भी कवर किए गए
  • मौजूदा validation मुख्यतः GS dump playback पर निर्भर है; PCSX2 hacking patch और mkfifo real-time test मौजूद हैं, लेकिन असली users तक पहुंचने के लिए emulator integration जरूरी है

paraLLEl-GS का लक्ष्य और शुरुआती बिंदु

  • paraLLEl-GS PlayStation 2 GS(Graphics Synthesizer) को Vulkan compute से emulate करने वाला project है
  • इसी author का 2020 का काम paraLLEl-RDP, N64 RDP को Vulkan compute में implement करता था, और Angrylion को baseline मानकर bit accuracy के करीब results और upscaling को लक्ष्य बनाता था
  • PS2 में GSdx लगभग 20 साल तक de facto state-of-the-art implementation बना रहा
  • 2014 के आसपास OpenCL-based PS2 GS compute implementation की कोशिश हुई थी, लेकिन पूरी नहीं हुई और मौजूदा upstream repository में भी उसे ढूंढना मुश्किल है
  • PS2 में compute shader raster इस्तेमाल करने की वजह N64 की तुलना में कम मजबूत है
    • PCSX2 में अच्छी तरह optimized software renderer और अपेक्षाकृत robust graphics-based renderer मौजूद हैं
    • software renderer upscaling support नहीं करता
    • graphics-based renderer, खासकर upscaling में, कई bugs और glitches दिखाता है
  • paraLLEl-GS hardware के लिए bit accuracy से ज्यादा स्पष्ट accuracy problems से बचने पर focus करता है
    • GSdx software renderer भी hardware bit-accurate implementation नहीं लगता, इसलिए direct comparison-based testing जल्दी ही सीमा पर पहुंच जाती है

PS2 GS क्यों मुश्किल है

  • GS वर्ष 2000 के हिसाब से theoretically प्रति सेकंड 1 अरब से ज्यादा pixels process कर सकने वाले fillrate और bandwidth वाला device था
  • VRAM 4MiB की छोटी है, लेकिन इसे कई DMA engines के जरिए लगातार stream होने के लिए design किया गया था
  • pixel pipeline अपने आप में N64 RDP की तुलना में कुछ हद तक सरल है
    • single texture
    • single-cycle combiner
    • बहुत basic anti-aliasing
  • कई features सामान्य graphics API से implement करना कठिन बनाते हैं
    • 1.0 से ऊपर blending: PS1 से आया व्यवहार, जिसमें 0x80 को 1.0 जैसा treat किया जाता है और 0xff तक represent किया जा सकता है
    • destination alpha test: destination alpha को pseudo-stencil की तरह इस्तेमाल किया जा सकता है
    • conditional blending: alpha के अनुसार blending को conditionally बंद किया जा सकता है
    • alpha correction: alpha record करने से पहले MSB को OR करके इसे जबरन 1 के करीब बनाया जा सकता है
    • alpha test का partial discard: केवल color छोड़कर depth write बनाए रखने जैसा behavior संभव है
    • AA1: coverage-to-alpha तरीका है और per-pixel depth write control से जुड़ा है
    • 32-bit fixed-point Z: D32_UINT support technically मौजूद है, लेकिन actual use case अभी नहीं देखा गया
  • programmable blending न हो तो immediate-mode desktop GPU पर ROV या per-pixel barrier की जरूरत होती है, जिससे performance काफी खराब होती है
  • compute implementation अपना tile-based deferred renderer(TBDR) बनाकर इन constraints को bypass करता है

Raster rules, vertex queue, memory layout

  • GS के primitives clip space में अपेक्षाकृत सामान्य तरीके से दिए जाते हैं
    • VU1 transform और clipping करता है और GS को कई vertex attributes output करता है
  • coordinates और attributes का GS-specific format होता है
    • X/Y: 12.4 unsigned fixed-point
    • Z: 24-bit या 32-bit uint
    • FOG: 8-bit uint
    • RGBA: per-vertex lighting के लिए 8-bit values
    • STQ: perspective-correct texture coordinates
    • UV: perspective correction के बिना 12.4 fixed-point unnormalized coordinates
  • raster rules D3D9 style के करीब हैं
    • triangles आधुनिक GPU की तरह top-left raster rule इस्तेमाल करते हैं
    • pixel center D3D9 की तरह integer coordinates पर होता है
    • lines Bresenham algorithm इस्तेमाल करती हैं, इसलिए upscaling मुश्किल है और rect या parallelogram से approximate करना पड़ता है
    • points nearest pixel पर snap होते हैं
    • sprites दो coordinates वाला simple quad हैं
  • GS vertex queue OpenGL 1.0 के immediate mode जैसी है
    • RGBA, STQ और कई registers set किए जाते हैं, और XYZ register write vertex “kick” बनाता है
    • TRIANGLE_FAN भी support करता है
  • PS2 pixel coordinates page units में layout किए जाते हैं
    • एक page 8KiB का होता है
    • page 32 blocks में बंटता है
    • 32-bit RGBA के हिसाब से page 64×32 pixels का होता है, और 32 8×8 blocks Z-order में swizzle होते हैं
  • 24-bit color या 24-bit depth में render करते समय बचे हुए ऊपरी 8 bits में texture रखी जा सकती है
    • 8H, 4HL, 4HH formats 8-bit और 4-bit palettes के लिए useful हैं

Textures, CLUT, TEXFLUSH

  • GS texturing में modern API जैसी चीजें और अजीब विशेषताएं दोनों मिलती हैं
    • texel center modern API की तरह half-pixel पर होता है
    • subtexel precision 8-bit नहीं बल्कि 4-bit लगती है
    • bilinear filter सामान्य bilinear है, N64 के 3-point filter जैसी special structure नहीं है
  • special addressing modes implementation कठिन बनाते हैं
    • REGION_CLAMP texture atlas के अंदर arbitrary region पर clamp apply कर सकता है
    • REGION_REPEAT हर coordinate पर (u & MASK) | FIX जैसे bit operations apply कर सकता है, इसलिए implementation और कठिन है
  • mipmapping derivatives की जगह interpolated Q factor के log2 और scale coefficient से LOD calculate करता है
    • compute implementation में derivatives पर निर्भर न होना बड़ा फायदा है
    • यह तरीका anisotropic filtering जैसी features support नहीं कर सकता
  • CLUT current palette रखने वाला 1KiB cache है
    • इस्तेमाल करने के लिए VRAM से CLUT cache में explicit copy करनी होती है
    • 32-bit color के हिसाब से यह एक 256-color palette रख सकता है
    • 16bpp के हिसाब से यह 16-color palettes के 32 सेट रख सकता है
  • TEXFLUSH texture cache synchronization और invalidation के करीब explicit command है
    • implementation की शुरुआत में TEXFLUSH को hazard tracking का आधार बनाने की कोशिश की गई, लेकिन आखिरकार इसे ignore करना पड़ा
    • games के TEXFLUSH भूल जाने या बहुत ज्यादा call करने की समस्या थी
  • final implementation ने minimal caching approach चुना
    • cache नहीं है ऐसा मानकर hazards को direct track करता है
    • feedback loops के लिए अलग exception handling पर विचार करता है
    • GSdx भी इसी दिशा में दिखता है

Vulkan compute rendering pipeline

  • implementation pipeline हर stage के बीच synchronization को premise मानकर बनाई गई है
    • CPU की VRAM copy को GPU से synchronize करना
    • VRAM upload या local-to-local copy करना
    • VRAM से CLUT cache update करना
    • VRAM को VkImage में unswizzle करके direct sampling के योग्य बनाना
    • rendering करना
    • GPU VRAM copy को वापस CPU से synchronize करना
  • सामान्य game behavior इस pipeline से अच्छी तरह मेल खाता है
    • texture को VRAM में upload करना
    • palette को VRAM में upload करना
    • CLUT cache update करना
    • texture से draw करना
    • जरूरत हो तो VRAM से VkImage में unswizzle करना
    • primitive batches को render pass के रूप में बनाना
  • backward hazards न हों तो batching और synchronization delay संभव है
    • इस तरह के renderer में performance पाने के लिए batch बनाए रखना important है
  • प्रमुख hazard cases अलग से handle किए जाते हैं
    • जिस VRAM पर copy पहले ही लिख चुकी है, उसी पर फिर copy
    • sampled texture या CLUT ने जिस VRAM को read किया है, उस पर copy
    • rendered area को texture के रूप में sample करना
    • rendered VRAM पर copy

Page tracking और texture cache

  • GS emulation में सबसे मुश्किल हिस्सा VRAM के read-after-write और write-after-write hazards संभालना है
  • 4MiB VRAM को पहले page units में बांटा जाता है
    • page framebuffer और depth buffer की unit है, इसलिए सबसे meaningful tracking unit है
  • page level पर tracked state इस प्रकार है
    • pending frame buffer write
    • pending frame buffer read
  • texture और VRAM copy में 256-byte alignment होता है, इसलिए 32 blocks के लिए u32 bitmask इस्तेमाल किया जाता है
    • VRAM copy write
    • VRAM copy read
    • CLUT cache या VkImage के लिए pending read
    • किसी भी write से overwritten block
  • 24-bit color में render करते हुए ऊपरी 8 bits को texture के रूप में sample करने पर hazard नहीं हो सकता
    • इसके लिए framebuffer write mask और texture read mask को अलग-अलग track किया जाता है
  • हर page के पास linked VkImage list होती है
    • page texture invalidated होने पर image destroy कर दी जाती है और VRAM से फिर unswizzle करना पड़ता है
    • एक texture कई pages में फैल सकती है, और उनमें से सिर्फ एक भी overwritten हो तो texture invalidated होती है
  • simple और conservative tracking भर से PS2 games में काम नहीं चलता
    • 256-byte block level tracking और write/read mask consideration जरूरी है
  • POT textures और REGION_CLAMP इस्तेमाल न होने से false positive पैदा हो सकते हैं
    • उदाहरण के लिए 512×448 render target को 512×512 texture के रूप में set करने पर unused area hazard जैसा दिख सकता है
    • implementation उस “red zone” के potential hazard को ignore करने वाली workaround इस्तेमाल करता है

CLUT batching और texture unswizzle

  • texture uploads batch करने के लिए CLUT uploads को भी साथ batch करना होता है
  • implementation CLUT की 1024 copies को snapshot ring buffer के रूप में रखता है
    • एक workgroup updates पर iterate करके SSBO में record करता है
    • N64 RDP के TMEM update जैसा है, लेकिन CLUT update कहीं ज्यादा simple है
  • Vulkan में नया VkImage allocate किया जाता है, VkDeviceMemory से suballocate किया जाता है, फिर compute shader से unswizzle किया जाता है
  • Vulkan specialization constants से texture format और swizzle logic को specialize किया जाता है
  • REGION_REPEAT का special behavior भी unswizzle stage में handle किया जाता है
    • इसके बाद ubershader को इस case के लिए manual bilinear filtering पर विचार करने की जरूरत कम हो जाती है
  • render targets भी VRAM SSBO के जरिए texture तक round-trip करते हैं
    • render target को texture में direct forward करने की कोशिश में बहुत ज्यादा bugs और exceptions माने गए

Triangle setup, binning, ubershader

  • paraLLEl-GS, paraLLEl-RDP की तरह tile-based renderer है
  • binning से पहले triangle setup किया जाता है, और input तीन arrays में बंटता है
    • position
    • per-vertex attributes
    • per-primitive attributes
  • rasterizer barycentric-based है और Fabian Giesen के graphics pipeline posts व Pineda 1988 paper में समझाए गए parallel rasterization method से काफी प्रभावित है
  • actual PS2 GS DDA, यानी scanline rasterizer है, लेकिन GS DDA का bit-accurate description पता न होने के कारण barycentric method इस्तेमाल किया गया
  • wide line और sprite implementation के लिए parallelogram भी support किया गया है
  • inv_area custom fixed-point RCP से calculate होता है
    • standard GPU RCP implementation के हिसाब से consistency में कमजोर है और करीब 22.5-bit accuracy देता है, इसलिए इसे avoid किया गया
    • custom RCP लगभग 24.0-bit accuracy target करता है
  • binning आम तौर पर 32×32 pixel blocks इस्तेमाल करता है
    • render pass प्रति primitives की maximum संख्या u16 index के कारण 64k है
    • देखे गए प्रमुख render passes आम तौर पर 10k~30k primitives range में थे
  • PS2 GS में fillrate ऊंचा और per-pixel complexity कम होने से pure ubershader संभव है
    • N64 के उलट bindless का फायदा लिया जा सकता है, इसलिए texturing complexity भी कम होती है
  • ubershader early Z, deferred on-tile shading, lazy pixel shading इस्तेमाल करता है
    • actual shading केवल तब की जाती है जब pixel previous result पर निर्भर करता है
    • alpha test, color write mask, alpha blending आदि ऐसी dependencies बनाते हैं
    • final framebuffer color और depth SSBO में record किए जाते हैं, जिससे GPU bandwidth use कम होता है

Supersampling और upscaling artifacts कम करना

  • केवल single-sample rendering से इस renderer की utility पर्याप्त नहीं है
  • उदाहरण के लिए 8x SSAA में GPU पर VRAM के 10 versions रखे जाते हैं
    • single-sample VRAM 1 copy
    • single-sample VRAM की reference value 1 copy
    • 8 supersamples
  • rendering के समय अगर single-sample VRAM और reference match करते हैं, तो supersample version load किया जाता है
    • incremental rendering में यह important है
  • tile पूरा होने पर clustered subgroup operations से multisample resolve किया जाता है और supersample व single-sample copies record की जाती हैं
  • supersampling simple upscaling की तुलना में jaggies कम करता है और 3D elements व UI elements के resolution feel को ज्यादा consistent मिलाता है
  • sprite primitives हमेशा single-rate पर render होने चाहिए
    • ये अधिकतर UI या similar elements होते हैं, और upscale करने पर intended rect के बाहर sample कर सकते हैं या bilinear filtering से जरूरत से ज्यादा blur हो सकते हैं
  • कई UI normal triangles से draw होते हैं, इसलिए कुछ flat primitives में attribute interpolation को single-pixel coordinates तक घटाया जाता है
    • perspective इस्तेमाल हो तब भी अगर सभी vertices का Q और Z समान है, तो इसे flat UI primitive माना जाता है
    • false positives हो सकते हैं, लेकिन tested games में यह पर्याप्त रूप से अच्छा काम करता है

Game-wise results और कठिन cases

  • Tales of the Abyss में PCSX2 Vulkan backend upscaling पर glass bloom alignment mismatch और square pattern दिखते हैं
    • paraLLEl-GS के 8x SSAA में खराब upscaling की typical problems ज्यादा नजर नहीं आतीं
    • संबंधित screenshot में FSR1 post-processing upscale भी applied है
  • Final Fantasy X का UI native resolution और 4x upscale comparison में upscaling problems दिखाता है
    • MSAA snap trick artifacts से बचने में effective है
    • core principle यह है कि UI को nearest neighbor integer scale से ज्यादा upscale न किया जाए
  • MGS2 में PCSX2 पर high blending accuracy मांगने वाले cases हैं
    • PCSX2 programmable blending path में हर primitive पर barrier डालता है, जिससे performance काफी गिरती है
    • paraLLEl-GS हमेशा 100% blend accuracy पर काम करने वाली structure है, और RX 7600 पर 16x SSAA scene 25W और 17% GPU utilization के साथ दिखाया गया है
  • Valkyrie Profile 2 में अपने ही pixel के alpha को palette index के रूप में sample करने का case है
    • paraLLEl-GS इसे detect करके texture index को special value बनाता है और in-register framebuffer color को reference करता है
    • इस optimization से render pass barrier 500 से ज्यादा से घटकर 18 रह गए
  • MGS2 intro का camo effect framebuffer को texture के रूप में sample करता है, लेकिन pixel alignment नहीं मिलने वाले overlapping coordinates इस्तेमाल करता है
    • PCSX2 भी यहां barrier add नहीं करता दिखता है, और paraLLEl-GS भी इसी तरह handle करता है
  • Shadow of the Colossus मजबूत stress test की तरह काम करता है
    • PCSX2 maximum blend accuracy पर intro में 2x upscale भर से GPU 24 FPS तक गिर जाता है
    • paraLLEl-GS 8x SSAA पर भी performance ठीक रखता है, लेकिन उस scene में load बढ़ जाता है
    • इस case में bottleneck GPU से ज्यादा CPU की geometry processing में है

मौजूदा स्थिति और अगले कदम

  • अभी testing का practical तरीका GS dump इस्तेमाल करना है
  • PCSX2 से raw GS trace dump कर सकने वाला hack-patch मौजूद है
  • mkfifo के जरिए rough real-time testing भी संभव है
  • end users के लिए useful बनने के लिए किसी न किसी रूप में emulator integration जरूरी है
  • PS2 library बहुत बड़ी है, इसलिए अभी भी कई bugs छिपे होने की संभावना काफी है
  • standalone library nature के कारण पुराने style के rendering API की तरह इस्तेमाल होने वाले potential use cases भी हैं

1 टिप्पणियां

 
GN⁺ 2024-09-03
Hacker News की राय
  • इस लेख में GS acronym को कब expand किया गया है, यह खोजने में काफी देर लगानी पड़ेगी क्या—ऐसा लगा। विषय दिलचस्प है, लेकिन यहीं पर थोड़ा पीछे छूटता महसूस होता है

    • यह Graphics Synthesizer का संक्षेप है, और Sony ने PS2 के “GPU” को यही नाम दिया था
  • “प्रोग्रामेबल blending होने की दुआ करो” — 2000 के दशक की शुरुआत में pixel shader पहली बार सीखने के बाद से मैं “blending shader” के ज़रिए programmable blending की उम्मीद करता रहा हूँ
    संदर्भ के लिए, custom texture formats/compression, texture composition आदि में उपयोगी “texture shader” के ज़रिए programmable texture decoding भी चाहता था
    अजीब बात है कि GPU को programmable blending से पहले ray tracing मिल गई; पहला तो गर्मियों की रात के सपने जैसा था, जबकि दूसरा बस एक और fixed-function block को programmable block में बदलने जैसा लगा। texture shader का अब भी इंतज़ार है

    • PowerVR परिवार से विरासत में आए mobile GPU में ऐसा feature है
      https://medium.com/pocket-gems/programmable-blending-on-ios-...
      https://developer.apple.com/videos/play/tech-talks/605
    • आखिरी बार जब मैंने देखा था, mobile PowerVR GPU में programmable blending था, और असल में blending का वही एक तरीका था। PS Vita पर blend state बदलने में करीब 1ms लगता था, इसलिए यह खास अच्छा नहीं था
    • VK_EXT_fragment_shader_interlock एक तरह का programmable blending ही लगता है। DirectX तरफ के Raster-order-views भी वैसे ही हैं
      इसका एक अच्छा उदाहरण है: https://vulkan.org/user/pages/09.events/vulkanised-2024/vulk...
    • आजकल mesh shader, work graphs, CUDA, और general C++ shader भी हैं। OTOY अब rendering पूरी तरह compute से करता है
    • इनमें से ज़्यादातर चीज़ें Vulkan या DX12 में emulate की जा सकती हैं, ऐसा लगता है; दूसरे API के बारे में पक्का नहीं। हालांकि असली use case को लेकर जिज्ञासा है
      कुछ हद तक implementation संभव मानता हूँ, लेकिन अगर कोई convincing use case न हो तो implementation work को justify करना मुश्किल होगा
  • GS में मुझे सबसे पसंदीदा हिस्सा bus architecture का बेहिसाब scale था। कुल 2560-bit width और cache partitioning भी चतुराई से की गई थी
    PS3 कुछ मायनों में पीछे कदम जैसा लगा, खासकर blending के मामले में

  • सोच रहा हूँ कि यह approach Dolphin के ubershader से कैसे compare होती है

    • बुनियादी तौर पर, दोनों लगभग तुलना के लायक नहीं हैं। Dolphin का ubershader modern flexible hardware पर fixed-function blending/texturing की नकल करने का एक काम करता है
      सच कहें तो Dolphin ने जब इसे अपनाया, तब भी यह technique पहले से पुरानी थी। यह project rasterizer तक शामिल करने वाला पूरा renderer है, और लेख में दिखने की तरह blending के लिए ubershader भी शामिल है
      shader triangles draw नहीं करता, बल्कि triangle के अंदर हर point पर call होकर कुछ inputs लेता है और उस point का color तय करता है। यह कुछ-कुछ पूरे CPU emulator की तुलना सिर्फ ADD/MUL instructions implement करने से करने जैसा है
  • top-left raster का क्या मतलब है, जानना चाहूँगा