मिरर डाइमेंशन में दिखाई दी बिल्ली
(substack.com/lcamtuf)- ऊपर से देखने पर यह एक महिला की तस्वीर लगती है, लेकिन इमेज पर DCT लागू करने पर frequency domain में छिपी बिल्ली सामने आती है—यह उसी का एक प्रयोग है
- time·space domain और frequency domain एक reversible transform से जुड़े होते हैं, इसलिए वही इमेज अलग-अलग तरीकों से समझी जा सकती है
- बिल्ली की तस्वीर को DCT से बदले गए “noise” pattern को कम opacity पर मिलाने से, original image बनी रहती है और transform के नतीजे में बिल्ली का आकार रह जाता है
- composite image में resizing के बाद भी बिल्ली की जानकारी बची रहती है, लेकिन enlarge करने पर यह tiles की तरह repeat होती है और shrink करने पर कट जाती है
- JPEG quality कम करने पर खासकर high-frequency components जोरदार quantize होते हैं, जिससे visually देखा जा सकता है कि lossy compression कितनी जानकारी फेंक देता है
DCT से बना “mirror dimension” प्रयोग
- frequency domain रोज़मर्रा के signals को उनके component waveforms के amplitudes में बदलकर interpret करने का तरीका है
- सबसे आम basis बढ़ती frequency वाली sine waves होती हैं
- दूसरे waveforms का इस्तेमाल भी alternative frequency domains बनाने के लिए किया जा सकता है
- इस transform की दो properties होती हैं
- original time domain या space domain data में वापस लौट सकने वाली reversibility
- एक ही mathematical operation से दोनों दिशाओं में transform करने वाली input-output symmetry
- compression में यह फर्क अहम है
- image को frequency domain में बदलने के बाद high-frequency components की precision घटाने या उन्हें हटाने पर भी result image perceptually मिलती-जुलती दिख सकती है
- इसी अनुपात में transmit या store करने वाला data कम हो जाता है
बिल्ली को frequency domain में छिपाने की प्रक्रिया
- शुरुआत बिल्ली की तस्वीर को Discrete Cosine Transform (DCT) से frequency-domain रूप में बदलने से होती है
- पिछले example की महिला वाली तस्वीर के ऊपर frequency domain का “cat noise” pattern कम opacity पर composite किया जाता है
- composite operation lossy होता है
- अपेक्षित result यह था कि महिला की तस्वीर DCT में अपेक्षाकृत uniform noise में टूटेगी, और insert किया गया cat noise फिर से cat image में इकट्ठा हो जाएगा
- वास्तव में composite image पर DCT लागू करने पर बिल्ली का आकार दिखाई देता है
- सीधे verify करने के लिए composite image और MATLAB example उपलब्ध हैं
- woman-with-cat.png
dct2(woman)से DCT calculate कर,imgaussfilt(cat, 1)से smooth किए गए result को display किया गया है
- resizing के बाद भी बिल्ली रहती है
- enlarge करने पर image tiles की तरह repeat होती है
- shrink करने पर image कट जाती है
- JPEG quality setting कम करने पर cat information खराब हो जाती है
- high JPEG quality पर image काफी ठीक दिखती है
- low quality पर high-frequency components से संबंधित bottom-right quadrant बहुत ज्यादा quantize हो जाता है
- यह visualization दिखाता है कि JPEG algorithm उन तरीकों से बहुत-सी जानकारी नष्ट करता है जिन्हें हम आसानी से नोटिस नहीं कर पाते
- audio spectrograms में hidden messages डालने के examples या JPEG DCT coefficients में text steganography जोड़ने जैसी चर्चाएं पहले से मौजूद हैं
- इस technique की practicality या पूरी novelty के बजाय, focus इस बात पर है कि frequency domain और time domain दिलचस्प तरीके से जुड़े हुए हैं
- bonus के तौर पर JPEG quality settings के हिसाब से degrade होती “standalone” frequency-domain cat की video है: https://vimeo.com/940487310/8a929a5eb5
1 टिप्पणियां
Hacker News की राय
पहचान में आने वाले subject वाली ज़्यादातर तस्वीरों में, यहाँ की तरह spectral energy origin, यानी ऊपर-बाएँ कोने के आसपास जमा होती है
https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_pr...
महिला की image का DCT भी ऐसा ही है। दूसरी ओर, photo का subject आम तौर पर frame के center की तरफ रखा होता है। इसलिए composite image में spatial-domain data और frequency-domain data एक-दूसरे में कम interfere करते हैं, और inverse transform करने पर बिल्ली का expression बचा रहता है
https://substackcdn.com/image/fetch/w_1456,c_limit,f_webp,q_...
महिला वाली image के लिए भी उल्टा यही principle लागू होता है
मैं check करना चाहता हूँ कि इस process को ऐसे समझना सही है या नहीं: a) महिला और बिल्ली की photo लेते हैं, b) बिल्ली को DCT से frequency domain में transform करते हैं, c) frequency-domain वाली बिल्ली को महिला की visual image पर composite करते हैं, d) composite image का DCT करने पर बिल्ली फिर से निकल आती है
और ठीक-ठीक कहें तो visual cat और frequency-domain वाली महिला का composite result आता है, लेकिन लगता है मतलब यह है कि visual cat ज्यादा prominent होती है
बहुत पहले एक student project में देखा था, जहाँ तक याद है, यह technique image हो या audio, किसी भी signal पर लागू होने वाली robust digital watermarking का आधार है
मुख्य use case यह है कि signal पर काफी processing होने के बाद भी copyright वाली material detect की जा सके। जैसे ripped या camcorder से shoot की गई movies, JPEG-2000 में दी गई material वगैरह
अगर movie industry में कोई व्यक्ति और technical details बता सके तो सुनना चाहूँगा
शायद Digimarc था; सोच रहा हूँ क्या वह Fourier transform based algorithm था
यह Fourier transform की time-frequency duality, यहाँ spatial-frequency duality, का बहुत अच्छा example है
Fourier transform का math इस बात की परवाह नहीं करता कि आप किस “direction” में transform कर रहे हैं, इसलिए time/frequency में मिलते-जुलते दिखने वाले functions के Fourier transforms भी frequency/time space में मिलते-जुलते होते हैं
इस case में, अगर बिल्ली के frequency plot को महिला के spatial plot में embed करें, तो महिला image के Fourier transform में बिल्ली दिखाई देती है, और इसका उल्टा भी सच है
यह steganography का बहुत cool और interesting application है
अगर आप illegal images को किसी सामान्य image के अंदर छिपाना चाहते हैं, तो उसे frequency domain में transform करके किसी दूसरी image में composite कर दें। receiver को बस reverse करने का तरीका पता हो, तो यह एक covert image transmission method बन सकता है जिसे detect करना मुश्किल हो
सामने वाले को पता न हो कि क्या ढूँढना है तो detect करना मुश्किल होगा, लेकिन पता हो तो आसान होगा
अगर hidden image को one-time pad के साथ combine करें, तो क्या उसे noise से अलग पहचान पाना असंभव नहीं होना चाहिए? lossy compressed images में noise naturally expected भी होता है। सोचता हूँ किसी ने यह पहले किया है या नहीं; और जब तक कोई बताए नहीं, शायद कभी पता भी नहीं चलेगा
Aphex Twin वगैरह ने भी ऐसा ही मजेदार trick इस्तेमाल किया था, जिसमें एक track के audio spectrogram में अजीब-सा चेहरा दिखता है: https://news.ycombinator.com/item?id=8509105
MetaSynth 90s के late दौर से मौजूद है, और audio के time-domain samples और frequency-domain images को transform करके Photoshop-style image filters के साथ combine करता है
https://uisoftware.com/metasynth/
लेख अंत तक जाते-जाते और बेहतर होता जाता है
भरोसा नहीं होता कि मुझे अभी जाकर समझ आया कि frequency domain को image compression में इस्तेमाल किया जा सकता है। देखने के बाद यह बहुत obvious लगता है। क्या ज़्यादातर image compression algorithms इसी तरह काम करते हैं? क्या वे frequency domain में weaker parts को बस हटा देते हैं?
हाँ। MP3, Ogg-Vorbis, JPEG सब इसी तरह काम करते हैं
कौन-सी frequencies रखनी हैं इसके लिए weights चुनना शायद psychoacoustic model पर based होता है, लेकिन मोटे तौर पर सचमुच high-order frequency information फेंक दी जाती है
DCT को अक्सर ज्यादा complex image या video compression algorithms के sub-step के रूप में भी इस्तेमाल किया जाता है
उदाहरण के लिए पहले image के ज्यादा-detail वाले कुछ regions ढूँढते हैं, उन regions पर DCT apply करके ज्यादा spectrum रखते हैं, फिर दूसरे regions पर भी यही करते हुए कहीं ज्यादा तो कहीं कम spectrum रखते हैं। video compression algorithms में जो quantization parameter देखा होगा, वही इस behavior को affect करता है
आम तौर पर high frequencies को पूरी तरह हटाया नहीं जाता, बल्कि उन्हें कम bits में encode किया जाता है
images असल मायने में band-limited नहीं होतीं, इसलिए सिर्फ frequency domain से उन्हें perfectly represent नहीं किया जा सकता
इसलिए compromise यह है कि उन्हें छोटे blocks में बाँटकर frequency domain और spatial-domain predictors को mix करके encode किया जाता है। फिर भी core idea broadly सही है
ज़्यादातर problem sharp edges से आती है। ऐसे edges को represent करने के लिए अनंत frequencies चाहिए, इसलिए उनमें से कुछ हटाने पर blur या ringing artifacts आते हैं
एक और वजह यह है कि band-limited signal अनंत तक repeat होता है, लेकिन real-world images ऐसी नहीं होतीं। photo के बाएँ तरफ जो है, वह ज़रूरी नहीं कि दाएँ तरफ क्या होगा इसका prediction करे
इसमें और भी factors हैं। सिर्फ frequencies फेंकना ही नहीं, बल्कि low-variance data को ज्यादा efficiently encode किया जा सकता है, यह भी important है
बात सिर्फ यह नहीं कि high-frequency information noise है, बल्कि आम तौर पर उसका magnitude भी छोटा होता है
JPEG 2000 और भी अजीब है। यह wavelet transform इस्तेमाल करता है
अगर JPEG 2000 file को बीच से काट दें, तब भी lower-resolution image recover की जा सकती है। किसी file length से नीचे जाते ही color information गायब हो जाती है और image grayscale बन जाती है
अगर बिल्ली ऊपर-बाएँ ज्यादा केंद्रित होती, तो शायद यह demo इतना अच्छा काम नहीं करता
DCT में बड़े low-frequency components काफी बनते हैं, और अगर बिल्ली ऊपर-बाएँ के पास होती तो वे components बिल्ली को ढक देते
position और frequency की quantum-mechanical representation में, यानी hbar को ध्यान में रखें तो position और momentum की representations में, इस तरह दो अलग-अलग functions को एक में ठूँसा नहीं जा सकता
क्योंकि जो functions सिर्फ position-dependent phase में अलग हों, वे भी अलग quantum states होते हैं