<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>XiangSiqi’s Blog</title>
        <link>https://blog.xiangsiqi.site/</link>
        <description>密涅瓦的猫头鹰在黄昏起飞</description>
        <lastBuildDate>Thu, 08 Oct 2026 04:17:16 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>zh-CN</language>
        <copyright>All rights reserved 2026, 向思齐</copyright>
        <item>
            <title><![CDATA[CS336: Basics]]></title>
            <link>https://blog.xiangsiqi.site/notes/336Basics</link>
            <guid>https://blog.xiangsiqi.site/notes/336Basics</guid>
            <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3df3936163f4801fb9bddf8cb1a08638"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-3df3936163f4802993aecc0f79df0f4a" data-id="3df3936163f4802993aecc0f79df0f4a"><span><div id="3df3936163f4802993aecc0f79df0f4a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3df3936163f4802993aecc0f79df0f4a" title="Tokenization"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Tokenization</span></span></h2><div class="notion-text notion-block-3df3936163f4802987e8c49477afd5b7">A Tokenizer:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3df3936163f4801aba39dd6936d1366b">For example:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3df3936163f480c2b6c5f9e12bc0a188"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Accb30cab-380d-4edf-b141-d55f50860836%3Aimage.png?table=block&amp;id=3df39361-63f4-80c2-b6c5-f9e12bc0a188&amp;t=3df39361-63f4-80c2-b6c5-f9e12bc0a188" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e03936163f48046bfe6fad5ce60fbee">The compression ratio is number of bytes per token, one could increase compression ratio by increasing <b>vocabulary size</b> (number of possible token values increases), leading to sparsity.</div><div class="notion-text notion-block-3e03936163f4807b93c0f6675fab82fe">There are approximately 150K Unicode characters, which means if you use every Unicode characters as a token(this is called character_tokenizer), you will have to face to problems:</div><div class="notion-text notion-block-3e03936163f480eba356e513e4300e02">Problem 1: this is a very large vocabulary.
Problem 2: many characters are quite rare (e.g., 🌍), which is inefficient use of the vocabulary.</div><div class="notion-text notion-block-3e03936163f480eaba87ebda8d83cc1d">Another choice is to use the result of Unicode-8:</div><div class="notion-text notion-block-3e03936163f480ef9d0ad8b4554d2ecb">this is byte tokenizer, the vocabulary is small however the compression ratio is one.</div><div class="notion-text notion-block-3e03936163f480ea856ac576cb76210e">Another ways is to split strings into words:</div><div class="notion-text notion-block-3e03936163f4805e96aeffe6ce62d5b3">What&#x27;s good: each token is meaningful (since humans invented words).
vocabulary_size = &quot;Number of distinct chunks in the training data&quot;</div><div class="notion-text notion-block-3e03936163f480638b5ac3fb10e8927b">Compression ratio is good, but vocabulary size can be huge.</div><div class="notion-text notion-block-3e03936163f480569e0cf8ca98325b66">Moreover:</div><ul class="notion-list notion-list-disc notion-block-3e03936163f480b8a286f2d684faf93f"><li>Many words are rare and the model won&#x27;t learn much about them.</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f4804490c8e2374ebce10a"><li>This doesn&#x27;t obviously provide a fixed vocabulary size.</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f480be82feea792e27e621"><li>New words we haven&#x27;t seen during training get a special UNK token, which is ugly and can mess up perplexity calculations.</li></ul><div class="notion-text notion-block-3e03936163f480be90d7ec212f4a8767">Popular tokenizer: Byte-Pair Encoding (BPE)</div><div class="notion-row"><a class="notion-bookmark notion-block-3e03936163f480a988b0d541f023484c" href="https://arxiv.org/abs/1508.07909" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Neural Machine Translation of Rare Words with Subword Units</div><div class="notion-bookmark-description">Neural machine translation (NMT) models typically operate with a fixed vocabulary, but translation is an open-vocabulary problem. Previous work addresses the translation of out-of-vocabulary words...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=3e039361-63f4-80a9-88b0-d541f023484c&amp;t=3e039361-63f4-80a9-88b0-d541f023484c" alt="Neural Machine Translation of Rare Words with Subword Units" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/1508.07909</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=3e039361-63f4-80a9-88b0-d541f023484c&amp;t=3e039361-63f4-80a9-88b0-d541f023484c" alt="Neural Machine Translation of Rare Words with Subword Units" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-3e03936163f4800b9feff0514ef1e5d3">Basic idea: <em>train</em> the tokenizer on raw text to construct a vocabulary tailored to the data.</div><div class="notion-text notion-block-3e03936163f480b3bf53e7c5f9afbaa4">Intuition: common sequences of bytes are represented by a single token, rare sequences are represented by many tokens.</div><div class="notion-text notion-block-3e03936163f480e887e7e880a669390b">Sketch: start with each byte as a token, and successively merge the most common pair of adjacent tokens.</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-3e03936163f480789658fcab52710aa6" data-id="3e03936163f480789658fcab52710aa6"><span><div id="3e03936163f480789658fcab52710aa6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e03936163f480789658fcab52710aa6" title="Resource Accounting"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Resource Accounting</span></span></h2><div class="notion-text notion-block-3e03936163f480c2b02ad9ac915c2b94">Question: How long would it take to train a 70B parameter model on 15T tokens on 1024 H100s?</div><div class="notion-text notion-block-3e03936163f48038977ef9b05415f6e6">According to</div><div class="notion-row"><a class="notion-bookmark notion-block-3e03936163f4805fb53ec127cdbb4928" href="https://arxiv.org/abs/2001.08361?utm_source=chatgpt.com" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Scaling Laws for Neural Language Models</div><div class="notion-bookmark-description">We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training,...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=3e039361-63f4-805f-b53e-c127cdbb4928&amp;t=3e039361-63f4-805f-b53e-c127cdbb4928" alt="Scaling Laws for Neural Language Models" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2001.08361?utm_source=chatgpt.com</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=3e039361-63f4-805f-b53e-c127cdbb4928&amp;t=3e039361-63f4-805f-b53e-c127cdbb4928" alt="Scaling Laws for Neural Language Models" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-3e03936163f480c68a06f03a3bbdea90"><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，so</div><div class="notion-text notion-block-3e03936163f48088a742dd9a7b69295c"><b>Question</b>: What&#x27;s the largest model that can you can train on 8 H100s using AdamW?</div><div class="notion-text notion-block-3e03936163f480a1bc20d1d8c1e3d4b5">Mixed precision training:</div><ul class="notion-list notion-list-disc notion-block-3e03936163f4803db347ce367edb8dfa"><li>Training with fp32 works, but requires lots of memory.</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f480c78a89edae2caee72e"><li>Training with fp16 and even bf16 is risky, and you can get instability.</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f480f0ad5cc934d39f741c"><li>Use bf16 for parameters, activations, and gradients</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f48080804bcd9fb32886a2"><li>Use fp32 for optimizer states(also for exponents)</li></ul><div class="notion-text notion-block-3e03936163f48071b49ef93668672e0e">By default, tensors are stored into memory of CPUs, so if you want to calculate it on GPUs, you need to move the sensor from CPUs to GPUs and move it back after calculation has been done.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e03936163f480e49b7df85c52bb7993"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Affc8f437-93fb-4ec0-8987-3d5d27f0e9f7%3Aimage.png?table=block&amp;id=3e039361-63f4-80e4-9b7d-f85c52bb7993&amp;t=3e039361-63f4-80e4-9b7d-f85c52bb7993" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e03936163f480c2a634eba9aad52713">You can also create tensors on GPUs.</div><div class="notion-blank notion-block-3e03936163f4808f9e5fdf328a3efbb4"> </div><div class="notion-text notion-block-3e03936163f4802db973d0cde30870c5">Then we need to talk about calculation of tensor, we usually use einops</div><div class="notion-text notion-block-3e03936163f48067ace5c0d07198a233">sum:</div><div class="notion-text notion-block-3e03936163f480b5a99efee0b5550e36">reduce:</div><div class="notion-text notion-block-3e03936163f48082b014d4e022606452">rearrange:</div><div class="notion-blank notion-block-3e03936163f480d19ffcd3fa4bedfc31"> </div><div class="notion-text notion-block-3e03936163f48002870fde8b1f8a4f2a">Then let’s talk about the computational costs, we usually use two notes:</div><ul class="notion-list notion-list-disc notion-block-3e03936163f480ab98bcca6dedac4ccc"><li>FLOPs: floating-point operations (measure of computation done)</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f4807a847bf8d674b2f336"><li>FLOP/s: floating-point operations per second (also written as FLOPS), which is used to measure the speed of hardware.</li></ul><div class="notion-text notion-block-3e03936163f4805e9b27dbf355b66ebc">For two matrix <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>， if we multiply them we need D times multiplication and D-1 times addition per output element, so we need:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e03936163f480edbb32ffcf92000a95">and for matrix <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>， if we do element-wise multiplication on them, we need <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> FLOPs.</div><div class="notion-blank notion-block-3e03936163f480e7914ff8981e404780"> </div><div class="notion-text notion-block-3e03936163f48036b416ce2e3cff364e">You can measure the practical FLOP/s and compare to ideal FLOP/s, then you can calculate Model FLOPs Utilization(MFU):</div><div class="notion-text notion-block-3e03936163f4808c8f5ac74127ca847b">Usually, MFU of ≥ 0.5 is quite good!</div><div class="notion-blank notion-block-3e03936163f4804495f5d47800c26350"> </div><div class="notion-text notion-block-3e03936163f4808f8d29c401f70932e7">So the important thing is that we need to figure out the bottleneck during computation, which means we need to know where the bottleneck coming from.</div><div class="notion-text notion-block-3e03936163f48069a6b3e5f2e37b62a0">Let’s begin with an example:</div><div class="notion-text notion-block-3e03936163f4802aa276e81dd8f4c168">here we need to move x from HBM to computation unit, and then we need to compute Relu and move y back to HBM:</div><div class="notion-text notion-block-3e03936163f480989229db221cae78bf">What is the bottleneck?</div><ul class="notion-list notion-list-disc notion-block-3e03936163f480ad834dcaff2e810108"><li>Memory-bound: communication time &gt; computation time</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f480779629dab7a82c6f05"><li>Compute-bound: computation time &gt; communication time</li></ul><div class="notion-text notion-block-3e03936163f4808ea278cadbdcb17372">So in this case, ReLU is memory-bound.</div><div class="notion-blank notion-block-3e03936163f4802d825bef7ddf39e510"> </div><div class="notion-text notion-block-3e03936163f480e18536e58814bba9ea">Or we can see accelerator intensity versus arithmetic intensity</div><div class="notion-text notion-block-3e03936163f48035a613dbf8ff5ce861">What is the bottleneck?</div><ul class="notion-list notion-list-disc notion-block-3e03936163f480a1a37bf8cae25abf88"><li>Memory-bound: arithmetic intensity &lt; accelerator intensity</li></ul><ul class="notion-list notion-list-disc notion-block-3e03936163f480e9853eec8579749c0e"><li>Compute-bound: arithmetic intensity &gt; accelerator intensity</li></ul><div class="notion-text notion-block-3e03936163f4809b9457efa0ccaade75">In general, we&#x27;ll find ourselves memory bound.</div><div class="notion-text notion-block-3e03936163f480b2a05de06f21eb0e79">And interesting is that if you change relu to geLu, gelu will do more flops than relu, so the arithmetic_intensity rise to 5, however when the problem is still memory bond so we won’t get faster or slower. So even the arithmetic intensity increase but time didn’t change the FLOPs/s will increase when facing memory bound, and when the bound type shift to compute bound the FLOPs/s will follow the speed of software.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e03936163f4802d9834fa348e8841c1"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:432px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ae0d9ad42-baf1-404b-9069-1e1a46387d3c%3Aimage.png?table=block&amp;id=3e039361-63f4-802d-9834-fa348e8841c1&amp;t=3e039361-63f4-802d-9834-fa348e8841c1" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e03936163f4800788f9c19370814977">For vector product the arithmetic intensity is 0.5, for matrix vector product is 1, however for matrix multiplication is n/3 where n is shape of matrix.</div><div class="notion-text notion-block-3e03936163f480cab71cfef5a56abd73">So when training for transformer what happens is compute-bound but for inference what mainly doing is matrix vector product, so it is compute-bond.</div><div class="notion-blank notion-block-3e03936163f480c99462ff115686f6b9"> </div><div class="notion-text notion-block-3e03936163f4801bae66deb08f3ca8b7">Now Let’s talk about a particular case:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e03936163f48082b0dec64f26648a0d"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A05503ab8-ca12-4516-9d12-898936443fe3%3Aimage.png?table=block&amp;id=3e039361-63f4-8082-b0de-c64f26648a0d&amp;t=3e039361-63f4-8082-b0de-c64f26648a0d" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e03936163f48009ae8dc196e98c6bf0">We have known that for forward propagation, a layer bring 2*B*D*D FLOPs, and for every weight matrix it brings 64 parameters, which means 4*64 bytes for float32.</div><div class="notion-text notion-block-3e03936163f480f58457d0d46faec62f">And for backward propagation, we need to compute gradients and store the tensor graph, usually every weight matrix and output have a same shape matrix.</div><div class="notion-text notion-block-3e03936163f480a890d2c7edd0846bb8">Let’s focus on one layer:</div><div class="notion-text notion-block-3e03936163f480c785ccd7940e8adda6">so we know that num_backward_flops = (2 * B * D * D) + (2 * B * D * D), which means for a full process of train we have to calculate 3 * 2 * B * D * D = 6 * B * (D^2), here B is the number of data and D^2 is number of parameters, this seems just work for easy MLPs but it turns out to be a good approximation for Transformers for short context lengths as well.</div><div class="notion-text notion-block-3e03936163f480eca277f049f58c3285">Another things is optimizer, for example AdaGrad, it needs to compute and store the square of grad as its optimizer state, so means to compute and store a matrix which has same shape of parameters:</div><div class="notion-text notion-block-3e23936163f48069b29cf0a5ea7ffc36">and for memory cost:</div><div class="notion-text notion-block-3e23936163f48091aaf1ca863d335911">here gradient has same shape as parameter for that when we use loss.backward the algorithm will release the non-leaf tensors automatically.</div><div class="notion-text notion-block-3e23936163f480c2880dddabc836b1a1">here activation memory is stored for backward propagation, and if you have enough memory, another way is to just store the value before activation, drop the activation and recompute them when doing backward propagation.</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-3e23936163f480a5b2c2fb8e154dc3d5" data-id="3e23936163f480a5b2c2fb8e154dc3d5"><span><div id="3e23936163f480a5b2c2fb8e154dc3d5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f480a5b2c2fb8e154dc3d5" title="Architecture"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Architecture</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f48033a8a3fb3275a9310d" data-id="3e23936163f48033a8a3fb3275a9310d"><span><div id="3e23936163f48033a8a3fb3275a9310d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48033a8a3fb3275a9310d" title="Normalization in or not in residual"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Normalization in or not in residual</span></span></h3><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4801aa7ffe0a53da764fd"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A20ff67a4-d539-45dc-9909-bb0443ef7613%3Aimage.png?table=block&amp;id=3e239361-63f4-801a-a7ff-e0a53da764fd&amp;t=3e239361-63f4-801a-a7ff-e0a53da764fd" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f48001827bf0a469ff83e7">Almost all modern transformers LLM use pre-norm but why?</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480668ef0f24ed14d4ffa"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Abf7ab316-b316-49bc-af98-80d2bf3fa0d8%3Aimage.png?table=block&amp;id=3e239361-63f4-8066-8ef0-f24ed14d4ffa&amp;t=3e239361-63f4-8066-8ef0-f24ed14d4ffa" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f4804bbbaddb3258d4f5e5">The pre-norm has better train stability for it has more stable gradients in deep networks:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480228cade268c132873e"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:384px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A954dddb0-40e6-4427-9f5f-47a59b8ba91e%3Aimage.png?table=block&amp;id=3e239361-63f4-8022-8cad-e268c132873e&amp;t=3e239361-63f4-8022-8cad-e268c132873e" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480d3a2e4f53fe6454536">for post-norm which means:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480adb38cc0108a669cf6">so that:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4809090c4eed06d0a8210">will cause numerical instability.</div><div class="notion-text notion-block-3e23936163f480f7a738ee2e136974de">So the key thing is whether put normalization in residual stream, we can also use post norm with normalization out of residual stream(Like Grok, Gemma 2). And even some models put normalization pre and post FFN and attention only if it is out of residual stream.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480328533de603660b5c7"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:336px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A17486142-d439-4272-bc60-19c16976d14a%3Aimage.png?table=block&amp;id=3e239361-63f4-8032-8533-de603660b5c7&amp;t=3e239361-63f4-8032-8533-de603660b5c7" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f4801ba225d0e3ab0b6910" data-id="3e23936163f4801ba225d0e3ab0b6910"><span><div id="3e23936163f4801ba225d0e3ab0b6910" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4801ba225d0e3ab0b6910" title="LayerNorm or RMSNorm"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">LayerNorm or RMSNorm</span></span></h3><div class="notion-text notion-block-3e23936163f48091b266fcb627fbd532">LayerNorm means:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480a28840dbfbb15d4382">RMSNorm means:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48052ad37f453e02b3e89">where:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4801e8e74c69b5a6955cd">There is not specific reason to use RMSNorm rather than LayerNorm, however, RMSNorm is usually faster and didn’t bring precision cost.</div><div class="notion-text notion-block-3e23936163f48015bd30d854f7e3ec3b">But it is interesting that in fact Layer-Norm just accounts for a insignificant FLOPs(typically <b>&lt; 0.5%</b>), but it accounts for a much more runtime(5% to 15%) because it is strictly <b>memory-bandwidth bound.</b> Its arithmetic intensity is around <b>0.5 to 1 FLOP/byte</b>.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4801c99c6f94fa6010d42"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A784f235f-511d-444e-bbb5-917d0083d5ef%3Aimage.png?table=block&amp;id=3e239361-63f4-801c-99c6-f94fa6010d42&amp;t=3e239361-63f4-801c-99c6-f94fa6010d42" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480c597dfc668d26ed718">More generally we always drop the bias in FFN.</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f480798d31f762aca86dc2" data-id="3e23936163f480798d31f762aca86dc2"><span><div id="3e23936163f480798d31f762aca86dc2" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f480798d31f762aca86dc2" title="Activations"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Activations</span></span></h3><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4804dbdb7de2c399d1fab"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A4fdeff5e-8945-4384-ae2c-ee02825aec8f%3Aimage.png?table=block&amp;id=3e239361-63f4-804d-bdb7-de2c399d1fab&amp;t=3e239361-63f4-804d-bdb7-de2c399d1fab" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480c4921fd3226dd5f4c1">In modern LLM, researchers have notice the importance of gate, so heuristically use it in activation:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4805bbe9fd973c7a70593"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A44de9c4b-2fae-4f8a-86d3-9014e9e87ffc%3Aimage.png?table=block&amp;id=3e239361-63f4-805b-be9f-d973c7a70593&amp;t=3e239361-63f4-805b-be9f-d973c7a70593" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f4809b87b3cadf129cc587">when we use GLU style activation, we have to store 3 matrix instead of 2 matrix so an idea is to use smaller ouput dimension by the factor 2/3 to keep the same number of parameters, but it is just a general rule instead of an iron rule.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4806ebae0ec6168417905"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:432px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A7a0a5e2f-186f-4207-88b9-2e3f4b84ab04%3Aimage.png?table=block&amp;id=3e239361-63f4-806e-bae0-ec6168417905&amp;t=3e239361-63f4-806e-bae0-ec6168417905" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f48090946fd3ef4e169c0a" data-id="3e23936163f48090946fd3ef4e169c0a"><span><div id="3e23936163f48090946fd3ef4e169c0a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48090946fd3ef4e169c0a" title="Serial versus Parallel"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Serial versus Parallel</span></span></h3><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480f3a997e4ff3a52d0fa"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Affe19216-668f-4458-acef-233c077e89bd%3Aimage.png?table=block&amp;id=3e239361-63f4-80f3-a997-e4ff3a52d0fa&amp;t=3e239361-63f4-80f3-a997-e4ff3a52d0fa" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f4806fbaeed3af16a36e34" data-id="3e23936163f4806fbaeed3af16a36e34"><span><div id="3e23936163f4806fbaeed3af16a36e34" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4806fbaeed3af16a36e34" title="Position Embedding"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Position Embedding</span></span></h3><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f48067b307d9174e398d10"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A0143f4f5-eb1d-471b-aa04-9b59d1de2663%3Aimage.png?table=block&amp;id=3e239361-63f4-8067-b307-d9174e398d10&amp;t=3e239361-63f4-8067-b307-d9174e398d10" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480e19aedf3989b4096dc">We have been familiar with the first two embeddings. Relative embedding doesn’t add PE to tokens vector instead it just directly participant into attention score calculation through <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> which is defined as:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480469daee34cda20f247">where k is a predefined maximum clipping distance.</div><div class="notion-text notion-block-3e23936163f48043b1ccce2975d001f2">But modern models usually use RoPE(rotary position embedding), the motivation of RoPE is that we want to find an embedding which makes the inner product of two embedded token vectors just has connection with their original vector and their relative position:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4803bbfa6c8c194be8cdb">Relative embeddings which directly modify attention matrix can’t fulfill this destination.</div><div class="notion-text notion-block-3e23936163f4802ab846d31e67a60501">Starting from two dimensional case, we have two original semantic vectors <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>, so we can rotate them by rotation matrix:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480e48064d6404489c364">so we can embed them:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480f99990dd4e031f449b">so we get:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480e7ae67c3e1b6876a7b">so the inner product is just about the original semantic vector and their relative distance.</div><div class="notion-text notion-block-3e23936163f4808aa8a0f0f9e2cdfa75">For high embedding dimension, we can just divide into several pairs and rotate the two dimensional vector of each pair separately, which means </div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f48010a74ee22d38b17e94"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A94bfb180-07d4-49bd-bc3b-348327faecbe%3Aimage.png?table=block&amp;id=3e239361-63f4-8010-a74e-e22d38b17e94&amp;t=3e239361-63f4-8010-a74e-e22d38b17e94" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f48053a2e6fdd415f84d39">in practice we write semantic vector <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> as <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> to compute, firstly we compute <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> for every pair:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480749744c0b6a55267dd">we share same <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> for same pair of different vector. then we compute rotation matrix element:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480cdb97bd386f0297acc">so:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480c29d47fa3f47c5f6be">where <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>. so:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4804f9c92ca0f21b8ede8">so the inner product of embedded vector is summarize of <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> of every pair.</div><div class="notion-text notion-block-3e23936163f480e9bdb6eb1b113aff1e">Here we use different angles for different pair for two reason:</div><ul class="notion-list notion-list-disc notion-block-3e23936163f480f98713fa6fa1b4aa0b"><li>Avoiding Periodic Collisions</li></ul><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48045abb0f11c92c1f112">Trigonometric function is periodic so it may cause aliasing between i-th token and i plus <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>-th token, different theta brings different frequency so it will never aliasing.</div><ul class="notion-list notion-list-disc notion-block-3e23936163f480ec8557d85d8c799049"><li>Having stronger expression</li></ul><div class="notion-text notion-block-3e23936163f4807e8400fb4827b25e4a">As we say Different theta brings different frequency, every pair are designed to express information of different fequency.</div><div class="notion-text notion-block-3e23936163f480dbb257f1d396d914f1">It is worth noting that absolute PE usually used in input layer while RoPE or other modern PE usually use in <b>every layer</b> of network.</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f48044a16fe64d11fb81e7" data-id="3e23936163f48044a16fe64d11fb81e7"><span><div id="3e23936163f48044a16fe64d11fb81e7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48044a16fe64d11fb81e7" title="Hyperparameters"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Hyperparameters</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3e23936163f4808eb15bd1f8fe45b400" data-id="3e23936163f4808eb15bd1f8fe45b400"><span><div id="3e23936163f4808eb15bd1f8fe45b400" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4808eb15bd1f8fe45b400" title="feed-forward ratio"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">feed-forward ratio</span></span></h4><div class="notion-text notion-block-3e23936163f480e79da0c5eeb24d9e49">Some consensus:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480a799fcf3e756a3da75">so for GLU, it turns to be:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4802fa872ed72650c20e8">feed-forward ratio(d_ff/d_model) is an typical hyperparameter</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f48040bb78fb92ea48730c"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:384px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A5fc1c63b-b0aa-4886-a4df-1b7be5ebe498%3Aimage.png?table=block&amp;id=3e239361-63f4-8040-bb78-fb92ea48730c&amp;t=3e239361-63f4-8040-bb78-fb92ea48730c" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f4809186c9df2bd579e8d7">there are a huge basin when use single-digit numbers of feed-forward ratio.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3e23936163f480748e2cd4a889a39a93" data-id="3e23936163f480748e2cd4a889a39a93"><span><div id="3e23936163f480748e2cd4a889a39a93" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f480748e2cd4a889a39a93" title="head dim ratio"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">head dim ratio</span></span></h4><div class="notion-text notion-block-3e23936163f480c2b692f2a1443ba2ca">we usually follow</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48094b7f1ee7e660562c3">which means the ratio between head dim and hidden dim is 1.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3e23936163f48042b03addbb3ad3f81e" data-id="3e23936163f48042b03addbb3ad3f81e"><span><div id="3e23936163f48042b03addbb3ad3f81e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48042b03addbb3ad3f81e" title="aspect ratio"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">aspect ratio</span></span></h4><div class="notion-text notion-block-3e23936163f4803aad06ea2ad2c2b01d">we define aspect ratio as ratio between dimension of model and number of layers.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f48080b54bf7907a729ab3"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:432px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A722c9855-b963-4fb9-9eae-b638177b724a%3Aimage.png?table=block&amp;id=3e239361-63f4-8080-b54b-f7907a729ab3&amp;t=3e239361-63f4-8080-b54b-f7907a729ab3" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480839380f1c00c8c9e61">however if the model going too deep, we will face the problem of parallelization, so we always want to go wide.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480ad91d0d7147e50bf13"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:432px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ab8c86223-a104-482d-bdb7-3284b0acc9d5%3Aimage.png?table=block&amp;id=3e239361-63f4-80ad-91d0-d7147e50bf13&amp;t=3e239361-63f4-80ad-91d0-d7147e50bf13" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f48057b548f2d2d39f062d">so aspect ratio is safe near 100.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3e23936163f4800383e9d273ea31f192" data-id="3e23936163f4800383e9d273ea31f192"><span><div id="3e23936163f4800383e9d273ea31f192" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4800383e9d273ea31f192" title="vocabulary sizes"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">vocabulary sizes</span></span></h4><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480c6bb62fd4c143937a6"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A701724c6-3c12-47ba-ad4d-020e941a258e%3Aimage.png?table=block&amp;id=3e239361-63f4-80c6-bb62-fd4c143937a6&amp;t=3e239361-63f4-80c6-bb62-fd4c143937a6" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f4805ebb38c649c3fd124d" data-id="3e23936163f4805ebb38c649c3fd124d"><span><div id="3e23936163f4805ebb38c649c3fd124d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4805ebb38c649c3fd124d" title="Regularization"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Regularization</span></span></h3><div class="notion-text notion-block-3e23936163f48059bd9ae45b6499698c">Nowadays we have tons of corpus on internet, so during pre-training we just even do a single pass on a corpus, which means overfitting is not a problem anymore, then what is the function of regularization(like weight decay or dropout)?</div><div class="notion-text notion-block-3e23936163f480168484eff5cdac9093">However weight decay is still a popular skill even when dropout seems to be mystifying, but the reason is not the rule of regularization which we measure by validation loss versus training loss:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480a4b8effaeeac5c8884"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:384px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A49921807-fb4d-4be3-ae77-937690dfe158%3Aimage.png?table=block&amp;id=3e239361-63f4-80a4-b8ef-faeeac5c8884&amp;t=3e239361-63f4-80a4-b8ef-faeeac5c8884" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480898d81f8c088e59b61">The rule of weight decay is that it brings progress when it interact with cosine decay optimization:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f48039a9b2da22e7b3dfed"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A99475b1e-6bf2-4ab2-91ee-0a05d58abecb%3Aimage.png?table=block&amp;id=3e239361-63f4-8039-a9b2-da22e7b3dfed&amp;t=3e239361-63f4-8039-a9b2-da22e7b3dfed" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f48017bc4adb67b0ba1fc1">the solid lines mean the primary training and dashed lines mean continued training at different checkpoint. In the two plot we can know that with cosine LR decay, weight decay could be helpful for training which is a little bit counterintuitive.</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f48027a242f830c25eb652" data-id="3e23936163f48027a242f830c25eb652"><span><div id="3e23936163f48027a242f830c25eb652" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48027a242f830c25eb652" title="Stability"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Stability</span></span></h3><div class="notion-text notion-block-3e23936163f48020a27fc86a39203036">You don’t want to get spikes in LLM training for that you usually can’t get fine final models when it happens. Here are some reason why you may get spikes or grad explosion.</div><ul class="notion-list notion-list-disc notion-block-3e23936163f48023887ff2696a0de02b"><li>Softmax, for its exponentials and dividing by zero.</li></ul><div class="notion-text notion-block-3e23936163f48002a2f7c4a135718c1b">the solution is to use z-loss regularization:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4800b8ed8cede1276735f">which means give penalty on z so that prevent training instability.</div><ul class="notion-list notion-list-disc notion-block-3e23936163f480b78f24c993dba9d2af"><li>QK norm</li></ul><div class="notion-text notion-block-3e23936163f4800d898cfd344d2f5eba">Adding normalization(RMS) on Q,K matrices is a key trick in keep training stable.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480f9a9a8c6a53e3886d2"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A806bc2c5-b73a-4e48-ab9f-cbae8d26d946%3Aimage.png?table=block&amp;id=3e239361-63f4-80f9-a9a8-c6a53e3886d2&amp;t=3e239361-63f4-80f9-a9a8-c6a53e3886d2" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f480f48472cd96eda7556b" data-id="3e23936163f480f48472cd96eda7556b"><span><div id="3e23936163f480f48472cd96eda7556b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f480f48472cd96eda7556b" title="Faster Decoding"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Faster Decoding</span></span></h3><div class="notion-text notion-block-3e23936163f480abbdc9f4de630bc598">A key problem happens in decoding. When using kv cache, we need to transport kV to calculation unit and compute logits of new token with new Q. The transportation of KV bring a very low arbitrary intensity. MQA and GQA was introduced to solve this problem.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480d2ba11f51355e6c144"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A53b1522b-4e20-4484-bbe2-1dac641bd97a%3Aimage.png?table=block&amp;id=3e239361-63f4-80d2-ba11-f51355e6c144&amp;t=3e239361-63f4-80d2-ba11-f51355e6c144" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480728d6df0abb6030db5">MQA(multi-query attention) use just one head for Key and Value matrices and broadcasting them to all heads when reasoning, meanwhile leading to lower expressiveness.</div><div class="notion-text notion-block-3e23936163f4801eb03ddd7aa9282663">So GQA(grouped-query attention) use multiple(but less than Query) heads for Key and Value matrices, and use key-query ratio to control the expressiveness.</div><div class="notion-text notion-block-3e23936163f480db8deefae1e1bb86ef">To turn a MHA model to a GQA model is easy, first thing is to do Heads pooling:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48031b2f6d93d08e9599d">and then do continuing pre-training (uptraining) on <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span><b> </b>of the original pre-training tokens.</div><div class="notion-row notion-block-3e23936163f480a5af26df2008778074"><div class="notion-column notion-block-3e23936163f480c1bae6f5c64ca3361e" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.375)"><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4800b9a94f1225fd7aa79"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:288px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A5dc0c896-0598-4e45-b68f-b59b82bcd2e2%3Aimage.png?table=block&amp;id=3e239361-63f4-800b-9a94-f1225fd7aa79&amp;t=3e239361-63f4-800b-9a94-f1225fd7aa79" alt="notion image" loading="lazy" decoding="async"/></div></figure></div><div class="notion-spacer"></div><div class="notion-column notion-block-3e23936163f480aaa19ce11063cdc249" style="width:calc((100% - (1 * min(32px, 4vw))) * 0.625)"><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f48009b84ee44fd2a650b5"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:336px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A2ac73521-fe1a-472a-a750-9d96c3b4f3a8%3Aimage.png?table=block&amp;id=3e239361-63f4-8009-b84e-e44fd2a650b5&amp;t=3e239361-63f4-8009-b84e-e44fd2a650b5" alt="notion image" loading="lazy" decoding="async"/></div></figure></div><div class="notion-spacer"></div></div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f4806090e0df36f7f617d0" data-id="3e23936163f4806090e0df36f7f617d0"><span><div id="3e23936163f4806090e0df36f7f617d0" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4806090e0df36f7f617d0" title="Hybrid Attention"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Hybrid Attention</span></span></h3><div class="notion-text notion-block-3e23936163f480c68e15c9bfbb8ae866">Another advantage pre-norm brings is that the output is simple addition of outputs of all layers:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480f2bd85c76ffb4c29ac">so that we can make some layer sparse(e.g. sliding window attention) or linear and so forth to accelerate computation.</div><div class="notion-text notion-block-3e23936163f4800eb5c5f7bc4e457993">However recent research has shown that too many non-standard attention layers will cause expressiveness degradation, so trade-off between them are important.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480a09ecde87b24aab095"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:384px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A5aa84bf7-0106-40cf-aa41-dfc535b20e0b%3Aimage.png?table=block&amp;id=3e239361-63f4-80a0-9ecd-e87b24aab095&amp;t=3e239361-63f4-80a0-9ecd-e87b24aab095" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480ff9307f4e72dbd83f3">This is what main researches nowadays work on.</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-3e23936163f480eca567e91513ece42a" data-id="3e23936163f480eca567e91513ece42a"><span><div id="3e23936163f480eca567e91513ece42a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f480eca567e91513ece42a" title="Attention Alternatives"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Attention Alternatives</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f48004ab24cfc78640d4b9" data-id="3e23936163f48004ab24cfc78640d4b9"><span><div id="3e23936163f48004ab24cfc78640d4b9" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48004ab24cfc78640d4b9" title="Linear Attention"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Linear Attention</span></span></h3><div class="notion-text notion-block-3e23936163f48001b392f046fdf4633c">The motivation is to remove the key non-linear part of standard attention: softmax</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480fca78ff740cf7a60b9">so that the computation complexity drop from <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> to <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>, the latter one is linear to <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>.</div><div class="notion-text notion-block-3e23936163f48020a2acef371b6a86af">One of the key reasons making linear attention an important work is that if you write it in incremental format it will be similar to recurrent neural network:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48021ad95e3f79a346653">here <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> is the memory at t-th step, it has fixed shape which means fixed memory volume.</div><div class="notion-text notion-block-3e23936163f480b58158dbb079c94046">But simple linear attention didn’t work well for strongly pooling and smoothing the memory <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3e23936163f48089b1bbe7688bef8c2b" data-id="3e23936163f48089b1bbe7688bef8c2b"><span><div id="3e23936163f48089b1bbe7688bef8c2b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48089b1bbe7688bef8c2b" title="Mamba"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Mamba</span></span></h4><div class="notion-text notion-block-3e23936163f480a09d96e1d2ec9b4132">Mamba introduce Gate mechanism into recurrent step of linear attention:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4804fa577dfd5af2bdee4">where <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> and <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> is the gate. Then we get:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4802fb7bfcebe85d90f82">which offer the way to be trained parallel.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3e23936163f4804fb6f5f35886455669" data-id="3e23936163f4804fb6f5f35886455669"><span><div id="3e23936163f4804fb6f5f35886455669" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4804fb6f5f35886455669" title="Gated delta net"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Gated delta net</span></span></h4><div class="notion-text notion-block-3e23936163f480ef81c5fa56510e2641">Gated delta net turn to be:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f480a2af8cfed4c53aea45">where <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>. So like LSTM, <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> is like forget gate and <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> is like update gate. And here <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> is meant to remove the similar information in memory <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> to <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>.</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e23936163f48044ba87edb64cb8a62c" data-id="3e23936163f48044ba87edb64cb8a62c"><span><div id="3e23936163f48044ba87edb64cb8a62c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f48044ba87edb64cb8a62c" title="Sparse Attention"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Sparse Attention</span></span></h3><div class="notion-text notion-block-3e23936163f480c7a766d42c93be8979">What Deepseek V3 stype sparse attention do is to filter a subset of tokens to compute. First is to compute weight of all tokens to current token:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48018957cd453c814890c">where</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48010ad63ef5a9f4846fa">and then filter tokens by Top-k</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f4803287ccff08f916331d">and compute attention on the subset:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e23936163f48004b30df76a20cf873c">It it worthing noting that DSA style sparse attention is not used for training, filter is training as a plugin after the model training.</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-3e23936163f4803bba50db8bd931c7a9" data-id="3e23936163f4803bba50db8bd931c7a9"><span><div id="3e23936163f4803bba50db8bd931c7a9" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e23936163f4803bba50db8bd931c7a9" title="Mixture of Experts"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Mixture of Experts</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e33936163f4807a8617c22f2f4a14eb" data-id="3e33936163f4807a8617c22f2f4a14eb"><span><div id="3e33936163f4807a8617c22f2f4a14eb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e33936163f4807a8617c22f2f4a14eb" title="Basics"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Basics</span></span></h3><div class="notion-text notion-block-3e23936163f48011a7cceb856ea3d2f7">MoE brings a lot of advantage through its sparsity, we can increase parameters of the model without increasing the FLOPs.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f4808b941ef6c51af536ea"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A0c4594a6-ab54-4b51-8095-4a6a2a4bb3e8%3Aimage.png?table=block&amp;id=3e239361-63f4-808b-941e-f6c51af536ea&amp;t=3e239361-63f4-808b-941e-f6c51af536ea" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f480e3b087fb979f16cc26">And MoE usually training faster:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e23936163f480018fb1cd8a1dd99bc6"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:432px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A914abbc6-fb64-4144-a6aa-98a4a6e009ca%3Aimage.png?table=block&amp;id=3e239361-63f4-8001-8fb1-cd8a1dd99bc6&amp;t=3e239361-63f4-8001-8fb1-cd8a1dd99bc6" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e23936163f48056a7b0f5a0f9b512ba">what we do is to construct multiple FFN for single layer:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e33936163f480aa88e9e470e803e29a">here t means t-th token, l means l-th layer, i means i-th FFN. And we need to decide which FFNs we want to use, we train a vector <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> for i-th FFN of l-th layer, and compute:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e33936163f4800fa43cf7ae59174822">as the score of i-th FFN, then get probability to choose i-th layer:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e33936163f480428b29ec275a7afb14">and then choose top-k FFNs and compute their output, and sum them with weight <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>:</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3e33936163f48009bb71cf53abf3344c">the second item is for residual connection.</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e33936163f480e2bbc1f70aeb71722b" data-id="3e33936163f480e2bbc1f70aeb71722b"><span><div id="3e33936163f480e2bbc1f70aeb71722b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e33936163f480e2bbc1f70aeb71722b" title="Shared MoEs"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Shared MoEs</span></span></h3><div class="notion-text notion-block-3e33936163f4803ba546d86bb5a67837">In shared MoEs, we maintain some FFNs to be used in every token:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e33936163f480049022ef3736208b75"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A2495930c-4386-4a21-8648-2cfd57c9bb8e%3Aimage.png?table=block&amp;id=3e339361-63f4-8004-9022-ef3736208b75&amp;t=3e339361-63f4-8004-9022-ef3736208b75" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e33936163f4803f82e9e643d0aed070">It’s easy to understand, and the effectiveness is significant.</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e33936163f4806289c5d1c6db0cff13"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Abc662b33-906b-49ed-b17d-133e555a4c00%3Aimage.png?table=block&amp;id=3e339361-63f4-8062-89c5-d1c6db0cff13&amp;t=3e339361-63f4-8062-89c5-d1c6db0cff13" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3e33936163f4804eabcbc05ac8e9bae5" data-id="3e33936163f4804eabcbc05ac8e9bae5"><span><div id="3e33936163f4804eabcbc05ac8e9bae5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e33936163f4804eabcbc05ac8e9bae5" title="Training MoEs"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Training MoEs</span></span></h3><div class="notion-text notion-block-3e33936163f4804cbfc2cc8fab84ba47">Major challenge in training MoEs is that if we want to sparsity, the gating decision are not differentiable.</div><div class="notion-text notion-block-3e33936163f48008a71ce36346078247">The popular way is shown below:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e33936163f48064bb1cf2fb161fc754"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Afaba9db1-679b-405d-a208-3a154a54768e%3Aimage.png?table=block&amp;id=3e339361-63f4-8064-bb1c-f2fb161fc754&amp;t=3e339361-63f4-8064-bb1c-f2fb161fc754" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3e33936163f48080b4c4dc1839045689">it is a heuristic algorithm, here P_i means the probability mass on expert i, and f_i means the numbers of token choosing expert i. When backward propagating, we stop the gradient of f_i, so when we penalize loss we actually penalize large P_i, which leading to sparsity. Here f_i is design to prevent cheating when router modulates the probability but still using just single expert.</div><div class="notion-text notion-block-3e33936163f480669397ddeec1ef25da">At perfect balance, <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>, then loss = alpha.</div><div class="notion-text notion-block-3e33936163f480e8a83be50e417daa23">Fine-tune on MoEs especially on experts will occur overfitting.</div><div class="notion-text notion-block-3e33936163f48060aa72ff5f6e535af9">Another way to train MoEs is to use upcycling:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e33936163f48043a5a7d32d163dd960"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Aaee022ed-4d8f-4e17-993b-2227665ce4e1%3Aimage.png?table=block&amp;id=3e339361-63f4-8043-a5a7-d32d163dd960&amp;t=3e339361-63f4-8043-a5a7-d32d163dd960" alt="notion image" loading="lazy" decoding="async"/></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3e33936163f4806bb5baec97a5f4956f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A2a765df0-3f9e-4137-bfa7-45f66e12ef6d%3Aimage.png?table=block&amp;id=3e339361-63f4-806b-b5ba-ec97a5f4956f&amp;t=3e339361-63f4-806b-b5ba-ec97a5f4956f" alt="notion image" loading="lazy" decoding="async"/></div></figure></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Chapter 5 Transformer]]></title>
            <link>https://blog.xiangsiqi.site/notes/DL5_cn</link>
            <guid>https://blog.xiangsiqi.site/notes/DL5_cn</guid>
            <pubDate>Mon, 27 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Attention Is All You Need]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-9b13936163f4831c82fc81f4af174484"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-b6c3936163f482519d2d817e52e05b99" data-id="b6c3936163f482519d2d817e52e05b99"><span><div id="b6c3936163f482519d2d817e52e05b99" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b6c3936163f482519d2d817e52e05b99" title="Transformer"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Transformer</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-95c3936163f482dca34d814bde2dc590" data-id="95c3936163f482dca34d814bde2dc590"><span><div id="95c3936163f482dca34d814bde2dc590" class="notion-header-anchor"></div><a class="notion-hash-link" href="#95c3936163f482dca34d814bde2dc590" title="CNN 与 RNN 的问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">CNN 与 RNN 的问题</span></span></h3><div class="notion-text notion-block-fe23936163f4827dad830121ea316612">CNN和RNN代表了这样一个深度学习时代，将人类的先验知识直接在模型中完全的表达出来。于是人类关于图像信息的平移不变性先验和在动物视觉研究得到的关于视觉皮层的理解被凝结到了CNN中，而关于序列信息的序列性先验和条件概率的思想被凝结到了RNN中。</div><div class="notion-text notion-block-f7f3936163f4831fb82981155e64f872">这种先验置入直接来讲是很有效的，CNN和RNN似乎准确表达了对于图像数据和序列数据的直接行为，但是先验同样也是限制。CNN的卷积核只能处理邻近卷积区域的内容，即使是深度CNN在结构上也是最终有限的。而RNN的序列结构对信息传播形成顺序强制，即使是Attention机制也只能在结构上缓解这种有限性。</div><div class="notion-text notion-block-ce13936163f483fa846681417d44c340">诚然CNN和RNN已经发展到这样一个状态，为了减弱CNN的问题产生了Deep CNN用来形成巨大的感受野和特征提取级次，由此派生出来Batch Normalization用以解决训练稳定问题，ShortCut用以解决退化问题；为了优化RNN产生了Encoder-Decoder用以职责解耦，产生了LSTM和GRU解决长程依赖问题，产生了Attention用以解决顺序强制问题。于是，当所有思路已经尽力后，CNN和RNN的的主要的问题就转向了自己，即架构本身。</div><div class="notion-text notion-block-d173936163f482e3a71b017b3048272a">于是，扬弃先验注入结构的思路，放弃CNN和RNN的结构，成为了最疯狂也是最合理的选择，Transformer机制横空出世。</div><div class="notion-row"><a class="notion-bookmark notion-block-4443936163f482398e75814826c2a7c0" href="https://arxiv.org/abs/1706.03762" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Attention Is All You Need</div><div class="notion-bookmark-description">The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=44439361-63f4-8239-8e75-814826c2a7c0&amp;t=44439361-63f4-8239-8e75-814826c2a7c0" alt="Attention Is All You Need" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/1706.03762</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=44439361-63f4-8239-8e75-814826c2a7c0&amp;t=44439361-63f4-8239-8e75-814826c2a7c0" alt="Attention Is All You Need" loading="lazy" decoding="async"/></div></a></div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3053936163f482d7875b019d828c3a1b" data-id="3053936163f482d7875b019d828c3a1b"><span><div id="3053936163f482d7875b019d828c3a1b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3053936163f482d7875b019d828c3a1b" title="注意力机制"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">注意力机制</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-25d3936163f4821e9d6d015f7ec2a18b" data-id="25d3936163f4821e9d6d015f7ec2a18b"><span><div id="25d3936163f4821e9d6d015f7ec2a18b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#25d3936163f4821e9d6d015f7ec2a18b" title="缩放点积注意力"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">缩放点积注意力</span></span></h4><div class="notion-text notion-block-75c3936163f4837f89930168830edb66">我们已经知道了在传统Attention机制中，attention指的是注意力权重<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，它表达着解码器中序列的第<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>项解码时对由BidiRNN编码的隐藏层<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>关注的程度，从而计算出上下文向量<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-f633936163f48280972081a0883aba90">因此<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>中同时编码了查询者<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>和被请求者<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>的<b>指称</b>和它们的关系，这在RNN中是合理的因为有稳定的序列结构，但是如果我们要抛弃RNN的结构，我们必须解耦这个结构为两个独立的可分离的向量，称为Query和Key向量，并且将隐层理解为数据本身，即Value。这就是QKV的出发点。</div><div class="notion-text notion-block-9d73936163f483e2a95381a84f899fcd">在经典 Attention 中，注意力分数通常由一个小神经网络计算出来，这叫<b>加性注意力</b>。它的形式是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，其中：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-4333936163f4830bbf0301934a847c9a">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是解码器当前状态，可以理解为“我现在想找什么信息”；<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是编码器第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置的隐藏状态，可以理解为“这个位置提供了什么信息”。为了判断二者是否匹配，模型先用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 把它们变换到同一个比较空间，再相加、过 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，最后用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 压成一个标量分数。</div><div class="notion-text notion-block-a463936163f4831ba22a013761ded854">所以加性注意力的本质是：给定两个隐藏状态，用一个小网络判断它们的相关程度。</div><div class="notion-text notion-block-3db3936163f482388ad481a97e95020f">Transformer 换了一种更清晰的拆分方式。它不再直接拿两个隐藏状态去做复杂比较，而是把每个 token 的表示投影成三种角色：</div><ul class="notion-list notion-list-disc notion-block-c7d3936163f482e1a2c581e1ceb5a6c1"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>（Query）：表示当前位置想要寻找什么信息。</li></ul><ul class="notion-list notion-list-disc notion-block-32c3936163f4832391c3816c110ee17d"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>（Key）：表示当前位置可以被怎样的查询匹配到。</li></ul><ul class="notion-list notion-list-disc notion-block-86c3936163f483958f2481a93f63f55a"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>（Value）：表示当前位置真正提供的信息内容。</li></ul><div class="notion-text notion-block-b4d3936163f482b9a9c501867673a850">这样一来，注意力计算就被分成两件事：先用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 计算“该看谁”，再用得到的权重去加权汇总 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中的信息。也就是说，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 主要负责匹配关系，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 才负责被传递的信息内容。</div><div class="notion-text notion-block-4463936163f48283b75901efe1c21fb3">因此，Transformer 中的注意力分数可以直接用点积表示：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9353936163f483918e06019e189add68">如果 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 方向相近，点积就大，说明第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置应该更多关注第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置；如果方向不相近，点积就小，注意力权重也会更低。</div><div class="notion-text notion-block-db93936163f48275a8fa01481e538923">写成矩阵形式，就是：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-6fb3936163f483e69b8901b76f5e0aaa">再用这个注意力矩阵加权汇总信息：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-bc53936163f482d49fb8011ae4747b17">这就是乘性注意力，也叫点积注意力。考虑防止 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的数值过大，还需要除以缩放因子 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-da73936163f48239919e01353521d6e7">注意到的是，此时 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的矩阵，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是序列长度，也就是 token 的个数。我们取 softmax 时是按 Query 的维度取的，即对同一行的所有项取 softmax。</div><div class="notion-text notion-block-c6c3936163f48209b72901503cab1661">更具体地说，假设输入序列有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token，每个 token 的表示维度是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，于是输入矩阵可以写成：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-5973936163f48369832281fab1450e51">通过三个线性变换得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-5da3936163f48356a0c381b4e2465010">因此：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-2913936163f4826a872881e961978116">这个矩阵就是注意力分数矩阵。第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 行第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 列的元素表示：第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token 作为 Query 时，对第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token 的 Key 有多关注。经过 softmax 后，第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 行所有数加起来等于 1，表示第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token 要从所有 token 处分别读取多少信息。</div><div class="notion-text notion-block-d373936163f483f990b001f76c97687e">最后再乘以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-4af3936163f482b59af701d68cf927d4">也就是说，输出仍然有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 行，每一行对应一个 token 的新表示，只是这个新表示已经融合了它从其他 token 读取的信息。</div><div class="notion-text notion-block-ec33936163f48359bb9881a7bd9d7045">举一个很小的例子。假设一句话有 3 个 token：</div><div class="notion-text notion-block-96e3936163f483f6989c8199c846ec0c">那么 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 矩阵：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-55a3936163f48214b8618123f0bf76b5">第一行表示“我”这个位置分别关注“我”“喜欢”“学习”多少；第二行表示“喜欢”这个位置分别关注三个词多少；第三行表示“学习”这个位置分别关注三个词多少。</div><div class="notion-text notion-block-9f03936163f48311a5b801f745d37f74">如果经过 softmax 后得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d1f3936163f4832eba0381ac984d6a9f">那么第二行 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的意思是：在更新“喜欢”这个 token 的表示时，模型从“我”读取 10% 的信息，从“喜欢”自己读取 30% 的信息，从“学习”读取 60% 的信息。这个例子不代表真实语言规律，只是说明注意力矩阵每一行的含义。</div><div class="notion-text notion-block-7823936163f48210af6a0123ed4b5609">再举一个维度例子。若一句话长度 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，模型维度 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，单头注意力里取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，那么：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-88b3936163f48356a99381c23508d851">于是：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-1e13936163f483539c9f01b405655322">这说明注意力矩阵描述的是 token 与 token 之间的关系，而输出向量描述的是每个 token 融合上下文后的新特征。</div><div class="notion-text notion-block-a143936163f483d8a7d6819991f8d2b6">总之这就是 Transformer 的 Scaled Dot-Product Attention。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-22d3936163f483eba29781f6c08bc12d"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:288px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A6a4ff6a6-c38b-4e10-85f9-00045faefe66%3Aimage.png?table=block&amp;id=22d39361-63f4-83eb-a297-81f6c08bc12d&amp;t=22d39361-63f4-83eb-a297-81f6c08bc12d" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-6523936163f482528ab581b15f625a2f">这里 Mask 是掩码矩阵，用于切断 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 之间的某些连接，强制让注意力不能看到某些位置。它最重要的用途之一是<b>自回归训练</b>。</div><div class="notion-text notion-block-a753936163f48288bd3001659f38bc05">在自回归语言模型中，模型要学习：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-1d73936163f482efbde00172129c2e17">也就是说，预测当前位置时，只能看到当前位置之前的 token，不能偷看未来 token。可是训练时为了并行计算，我们通常会把整句话一次性输入 Transformer。这样如果不加限制，第一个位置就可能看到第二个、第三个位置的信息，训练目标就被破坏了。</div><div class="notion-text notion-block-b373936163f48246acc281755ce2a6fc">因此需要使用 causal mask（因果掩码）。它的作用是：第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token 只能关注第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置及其之前的位置，不能关注未来位置。</div><div class="notion-text notion-block-e193936163f483b7bf598127e9fd440e">例如序列为：</div><div class="notion-text notion-block-4c73936163f48391930d013d5b707a71">如果按行表示 Query 位置，按列表示 Key 位置，那么不加 mask 时，每个 token 都可以看所有 token：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a453936163f483668c978168100562d7">这里的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示不屏蔽。为了做自回归训练，需要屏蔽未来位置，mask 矩阵可以写成：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a0b3936163f482edb9c3817c77dd0eb2">第一行表示“我”只能看“我”，不能看未来的“喜欢”和“学习”；第二行表示“喜欢”可以看“我”和“喜欢”，但不能看未来的“学习”；第三行表示“学习”可以看前面所有 token。</div><div class="notion-text notion-block-3ca3936163f4825e9cc98194ae61bcbe">计算时把这个 mask 加到注意力分数上：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9323936163f48387836181179cae961b">由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，被 mask 的位置在 softmax 后权重会变成 0，也就不会参与信息读取。这样，Transformer 虽然一次性并行处理整句话，但每个位置仍然只能使用它在自回归生成时应该能看到的信息。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-3f43936163f482cab6f68159d70c00a3" data-id="3f43936163f482cab6f68159d70c00a3"><span><div id="3f43936163f482cab6f68159d70c00a3" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3f43936163f482cab6f68159d70c00a3" title="多头注意力"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">多头注意力</span></span></h4><div class="notion-text notion-block-db83936163f48333adfd819a23243720">这种注意力机制能够提升效率和有效性的关键在于多头注意力。它的意思是：对于同一个输入，不只计算一次注意力，而是让模型从多个不同的子空间中分别计算注意力。每一个子空间由一个头 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 完成。</div><div class="notion-text notion-block-9093936163f482b1a2a3816f9fd830b8">这里直接从输入矩阵 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 出发：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d943936163f482cca4aa8119382f1eda">第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个注意力头有自己独立的三组投影矩阵：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-8743936163f4820c98b68132ae302be6">于是第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个头先构造自己的 Query、Key、Value：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d443936163f483a7add401b69334c82f">然后输入缩放点积注意力：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-1cf3936163f48207b707814746753611">每个头的输出维度为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-2b53936163f4827f9ca701ae16554097">如果一共有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个头，把它们在特征维度上拼接起来，就得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-cb13936163f482378a9781ef593cb7fa">然后再通过一个输出矩阵 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 混合：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-2e23936163f4832287670133f9a6d5e3">其中：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-6cc3936163f48376b4f581ff5e831a2b">因此最终输出维度为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-1153936163f48202aa5e81f2aba6354c">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 用来把多个头的信息重新整合回模型的表示空间。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-aeb3936163f483be9a7e817fcb984db9"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:288px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A6eb1a380-5971-4a75-8a27-e2c31a872244%3Aimage.png?table=block&amp;id=aeb39361-63f4-83be-9a7e-817fcb984db9&amp;t=aeb39361-63f4-83be-9a7e-817fcb984db9" alt="notion image" loading="lazy" decoding="async"/></div></figure><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-0593936163f48217bded01448453a358" data-id="0593936163f48217bded01448453a358"><span><div id="0593936163f48217bded01448453a358" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0593936163f48217bded01448453a358" title="交叉注意力与自注意力"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">交叉注意力与自注意力</span></span></h4><ul class="notion-list notion-list-disc notion-block-e1a3936163f48310ab1f81a671637892"><li>Cross Attention</li></ul><div class="notion-text notion-block-0583936163f482f09d1c817c974c0c7f">传统Attention是一种Cross Attention，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>的形成来源于Decoder，而<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>来源于Encoder。</div><div class="notion-text notion-block-e723936163f483158a680127e939d51a">因此：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-85b3936163f4834e9e0e01098ffc32f1">因此Cross Attention是一种高效的解码器，在Transformer的解码器结构中，核心模块就是交叉注意力。</div><div class="notion-text notion-block-2513936163f482f9ad93012de12a7147">以往的解码器，对于RNN without attention，输入会被在推进过程中被污染和破坏，造成长程关联弱，对于CNN，利用的是全连接层，是稠密的没有稀疏结构。</div><div class="notion-text notion-block-54a3936163f482eeaa048107507759e3">Cross Attention保证了信息的无损性，且不需要连续推进过程，一次完成，可以并行，注意力机制构造了FFN没有的稀疏指向结构。</div><ul class="notion-list notion-list-disc notion-block-e1a3936163f48213b4ea01247d347187"><li>Self Attention</li></ul><div class="notion-text notion-block-da73936163f4826197ea815bf9469f3d">只有当Q,V解耦，扬弃RNN后才会产生自注意力机制，对于输入<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，直接构造：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-5163936163f4824ca5a101a3b3cd4ccf">通常自注意力机制是作为一个高效的编码器存在的，例如在Transformer的编码器结构中，核心的模块就是自注意力机制。</div><div class="notion-text notion-block-2393936163f483bf869d019b992dce2d">以往的编码器，对于RNN，需要顺序的传递序列信息，通常造成长程关联弱，对于RNN，需要不断加深网络来建立大尺度连接，而对于Self Attention，信息的互访是一步完成的。</div><div class="notion-text notion-block-0e43936163f483a789bd0153a5a3dc8b">并且Self Attention机制的拥有极其领先的自指思想。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3c43936163f480e68f37d92911514266" data-id="3c43936163f480e68f37d92911514266"><span><div id="3c43936163f480e68f37d92911514266" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3c43936163f480e68f37d92911514266" title="Implementation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Implementation</span></span></h3><div class="notion-text notion-block-3c43936163f480428ddfc97950477a59">值得注意的是，代码中为了减少计算成本我们是一次性计算qkv：</div><div class="notion-text notion-block-3c43936163f480ae9a4cdf39c089b08d">再做多头分解：</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-5bf3936163f483f2b5e0810d83abc3b5" data-id="5bf3936163f483f2b5e0810d83abc3b5"><span><div id="5bf3936163f483f2b5e0810d83abc3b5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5bf3936163f483f2b5e0810d83abc3b5" title="Transformer in NLP"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Transformer in NLP</span></span></h2><div class="notion-text notion-block-0d03936163f482dfbc7101733ca51aa2">Transformer让自然语言处理进入一个新时代，从Transformer(2017)开始几乎所有的自然语言方向的架构都基于Transformer的变种。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-9493936163f4831cac3901e2b1dc738d" data-id="9493936163f4831cac3901e2b1dc738d"><span><div id="9493936163f4831cac3901e2b1dc738d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#9493936163f4831cac3901e2b1dc738d" title="Transformer 架构"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Transformer 架构</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-8053936163f4833abe8c816ea0d43c50" data-id="8053936163f4833abe8c816ea0d43c50"><span><div id="8053936163f4833abe8c816ea0d43c50" class="notion-header-anchor"></div><a class="notion-hash-link" href="#8053936163f4833abe8c816ea0d43c50" title="位置编码"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">位置编码</span></span></h4><div class="notion-text notion-block-f253936163f48240ab94810f8530cfd7">Transformer 的直接问题是，由于不采用 RNN 的递推结构，数据的序列性先验没有自然注入到模型中。对于自注意力层，我们有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-8643936163f483bdaca001490ba56e25">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是 token 数量，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的每一行表示一个 token 的向量。<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 作用在特征维度上，而不是作用在 token 位置上。</div><div class="notion-text notion-block-94e3936163f483cb8b99812b19ad01cd">更准确的说法是：不带位置编码的 self-attention 对 token 顺序具有<b>置换等变性</b>。设 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是一个置换矩阵，用来交换输入 token 的顺序，那么：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-5ec3936163f4822baaee81ab2ff9cdfe">于是：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-66f3936163f483b7af8c81c9df386d0c">同理：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-2a73936163f483c0953981bc2f66cec4">接下来考虑注意力矩阵。设原始注意力分数矩阵为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-7703936163f483c0826b8178103a7994">交换输入顺序后，新的注意力分数矩阵为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9043936163f4831ab35b8178d0f8463f">代入 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a873936163f482539e1681677d4e582b">由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-fc23936163f48380a877814093cc1758">这说明注意力矩阵并没有保持完全不变，而是行和列都按照同一个置换被重新排列了。直观地说，如果输入 token 的顺序被交换，那么“谁查询谁”的关系表也会跟着交换。</div><div class="notion-text notion-block-eb13936163f4823c8ab38107c97d9fd6">再看注意力输出。忽略缩放和 softmax 的细节，原始输出可以写成：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-e313936163f4839c892b8125ea8601fb">交换后有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-85a3936163f48346882b81977f7b5801">由于置换矩阵满足 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3f73936163f483dfaa35015ce8f6c6b0">也就是说，输入顺序被置换后，输出也只是按照同样方式被置换。交换并没有被“抵消”成原来的结果，而是一路传递到了输出中。</div><div class="notion-text notion-block-0863936163f483db908901f046af6a6a">这就是置换等变性的含义：如果输入顺序改变，输出顺序也随之改变。更直观地说，假设我们已经知道原输入 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的输出是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，那么当输入变成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，我们甚至不需要重新推理，也能知道输出一定是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，也就是把原输出按同样方式交换即可。</div><div class="notion-text notion-block-cda3936163f483fc9a7601309f2a635b">这说明不带位置编码的 self-attention 并没有真正利用 token 的绝对顺序。它会处理 token 之间的内容关系，但没有额外机制知道“这个 token 原本在第几个位置”。因此，如果任务需要区分不同词序，模型就必须获得某种位置信息。</div><div class="notion-text notion-block-c5f3936163f483b8ae0a81e982980a69">因此没有位置编码时，Transformer 无法区分同一组 token 的不同排列。对于序列文字“<em>技术之本质只是缓慢地进入白昼</em>”和“<em>白昼之本质只是缓慢地进入技术</em>”，如果不额外提供位置信息，模型只能看到相同 token 的集合，很难区分它们的先后结构。</div><div class="notion-text notion-block-8f83936163f482b9a8b401abc5bb2dbc">为了解决这个问题，必须对数据做一次处理，这个处理传入了数据的位置先验，称为Positional Encoding，考虑在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>加上<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>项以破坏交换对称性：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d933936163f483858150010af0a52f64">这样交换对称性就被破坏了。有许多方法构建<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>矩阵。</div><ul class="notion-list notion-list-disc notion-block-9a13936163f4829e8e23817f1ed91680"><li>Sinusoidal PE</li></ul><div class="notion-text notion-block-a733936163f48378a3de81bd71685ed9">这正是Transformer自己使用的思路：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9743936163f483f180a681a65bc275ee">这里<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>就是词的位置（索引），而<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>是特征向量（例如embedding）的维度索引，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>是特征向量的维度数，故<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-9233936163f483429743819dfa44810d">这个编码的优势在于它是平移变换幺正的：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-e933936163f4833c9fb601bddaf5e9f6">这一性质保证了，位置移动不会改变PE矩阵的范数，不会在过程中发生衰减或者爆炸，保证训练稳定性和保证能量与信息量的稳定性。并且模型有能力学会相对位置关系，即学习相对的相位。</div><div class="notion-text notion-block-c633936163f4827aa50381807cc8697a">最后我们考虑当<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>时保持不变的情况，这意味着编码相同，可以导出满足的条件为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-e963936163f483c8919901e7ae859c1f">这意味着只有在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>在前期可能有相同编码，并且由于证书限制，这个概率是更小的。</div><div class="notion-text notion-block-e8b3936163f482e5904e815d228e033a">但是这种编码的特性是加性编码，它将PE直接加到Embedding上去，这在一定程度上破坏了语义，后来发展起来的RoPE解决了这个问题，保留了幺正变换特性。</div><ul class="notion-list notion-list-disc notion-block-8493936163f483cf82cb01f297d93ab5"><li>Rotary Positional Encoding</li></ul><div class="notion-text notion-block-4db3936163f4839bb2ef81baafdb4b2e">为了不破坏Embedding，RoPE考虑的是对注意力矩阵做Rotate，考虑<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，此时<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>每一行代表一个position，因此取出第<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>行<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>和第<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>行<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，分别旋转<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>度和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>度：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-46b3936163f4827289fb81e4bef493a3">其中<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>是旋转矩阵：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-5923936163f48355bc520178fa55e9f2">可以看到如果<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，那么当<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>时就完成了一个周期，这在长文本是非常不利的，因此我们同样引入正余弦位置编码中的频率项：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><ul class="notion-list notion-list-disc notion-block-3663936163f4827aa8a881ca7ea6c998"><li>可学习PE：通过将PE当做一个linear层，缺陷是必须指定大小。</li></ul><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-e3e3936163f482c8853f0158e0c81678" data-id="e3e3936163f482c8853f0158e0c81678"><span><div id="e3e3936163f482c8853f0158e0c81678" class="notion-header-anchor"></div><a class="notion-hash-link" href="#e3e3936163f482c8853f0158e0c81678" title="架构"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">架构</span></span></h4><div class="notion-text notion-block-4ea3936163f482319a4481161118bb53">现在我们就容易得到整体的Transformer architecture:</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-ffe3936163f482eb90c9010b23aaf171"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:432px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A22eee89c-bf5a-4104-b3e6-88407472629e%3Aimage.png?table=block&amp;id=ffe39361-63f4-82eb-90c9-010b23aaf171&amp;t=ffe39361-63f4-82eb-90c9-010b23aaf171" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-eb93936163f482e1aaac81e948a35259">对于编码器部分，每一层主要由多头自注意力和前馈网络组成。多头自注意力负责让不同 token 之间交换信息，前馈网络负责对每个 token 自身的特征表示做非线性增强。每个子层都会配合残差连接和 LayerNorm，而不是 BatchNorm。</div><div class="notion-text notion-block-f5b3936163f4830793710196c716e015">对于解码器部分，每一层多了一个结构：先使用掩码多头自注意力，保证自回归生成时当前位置不能看到未来 token；然后使用交叉注意力，让解码器读取编码器输出；最后再经过前馈网络。每个子层同样配合残差连接和 LayerNorm。</div><div class="notion-text notion-block-1373936163f4835386af0131b6206ad4">我们最后需要关注的是 Feed Forward 层。它通常称为 <b>position-wise feed-forward network</b>，也就是说，它对每个 token 独立应用同一个 MLP。这里输入 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是一整段序列的表示矩阵，形状为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-7363936163f4836288998111c0ae3b0e">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是 token 数量，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是每个 token 的特征维度。<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的每一行代表一个 token，因此可以写成：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-4663936163f48383b08881bc332da017">通常有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-92a3936163f482aa80cf81e8c69c10a1">所以中间层形状为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-40c3936163f483e08ad081040a99ac07">最终输出又回到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-e963936163f483bcb15a014d1bc9a43e">等价地，对第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token 的表示 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-fc73936163f48289ad4e01f08deeb737">这说明 FFN 不负责 token 之间的信息融合；token 之间的信息融合已经由 attention 完成。FFN 的作用是在每个 token 自身的特征维度上做表达增强：先用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 投射到更高维度，经过非线性激活后，再用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 投射回 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-e6a3936163f483c1a75e01cc70b5c908">这个结构类似于 CNN 中的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 卷积：它不直接混合空间位置，而是在每个位置上独立加工通道特征。Transformer 的 FFN 也是类似的，它不直接混合 token 位置，而是在每个 token 上独立加工特征维度。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-f6b3936163f482a8b9d88107fc93931d" data-id="f6b3936163f482a8b9d88107fc93931d"><span><div id="f6b3936163f482a8b9d88107fc93931d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#f6b3936163f482a8b9d88107fc93931d" title="自回归语言建模与 GPT"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">自回归语言建模与 GPT</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-b033936163f482d9bf138145d39526b3" data-id="b033936163f482d9bf138145d39526b3"><span><div id="b033936163f482d9bf138145d39526b3" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b033936163f482d9bf138145d39526b3" title="GPT-1 与预训练微调"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">GPT-1 与预训练微调</span></span></h4><div class="notion-text notion-block-4ad3936163f483c4a4ac011bb24a0e54">OpenAI在2018年提出了GPT 1</div><div class="notion-row"><a class="notion-bookmark notion-block-0283936163f482618a51812fdf1e2ce2" href="https://openai.com/index/language-unsupervised/" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Improving language understanding with unsupervised learning</div><div class="notion-bookmark-description">We’ve obtained state-of-the-art results on a suite of diverse language tasks with a scalable, task-agnostic system, which we’re also releasing. Our approach is a combination of two existing ideas: transformers and unsupervised pre-training. These results provide a convincing example that pairing supervised learning methods with unsupervised pre-training works very well; this is an idea that many have explored in the past, and we hope our result motivates further research into applying this idea on larger and more diverse datasets.</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Fopenai.com%2Fapple-icon.png%3Fd110ffad1a87c75b?table=block&amp;id=02839361-63f4-8261-8a51-812fdf1e2ce2&amp;t=02839361-63f4-8261-8a51-812fdf1e2ce2" alt="Improving language understanding with unsupervised learning" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://openai.com/index/language-unsupervised/</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Fimages.ctfassets.net%2Fkftzwdyauwt9%2Fbbab3cf5-a807-49f0-9f93ff442ad7%2F2cb463c19ae4d17e661a8a18ce574f6b%2Fimage-19.webp%3Fw%3D1600%26h%3D900%26fit%3Dfill?table=block&amp;id=02839361-63f4-8261-8a51-812fdf1e2ce2&amp;t=02839361-63f4-8261-8a51-812fdf1e2ce2" alt="Improving language understanding with unsupervised learning" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-2c73936163f483ccbaf08145071a7e1d">传统的NLP有许多监督/自监督任务用于训练模型，哪种方法最优尚未达成共识。并且，通过这些方法学习到的表征如何运用到具体的任务也为形成统一。GPT 1的思路是，使用<b>自监督的预训练和监督微调</b>。</div><div class="notion-text notion-block-8293936163f48318a0030174a02e00dd">在预训练阶段，对海量文本进行顺序自回归语言建模，只使用transformer的decoder，并且带有掩码注意力：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-f4d3936163f483ac8b5201b188a524a1"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:240px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A75799693-457b-4e30-860f-32852679197e%3Aimage.png?table=block&amp;id=f4d39361-63f4-83ac-8b52-01b188a524a1&amp;t=f4d39361-63f4-83ac-8b52-01b188a524a1" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-f503936163f483e68b55019bc9a8885a">注意到，GPT将batch norm换成了layer norm，这是因为</div><div class="notion-text notion-block-2dc3936163f4833cb28a81f64f73a646">1.在大规模训练中文本长度是动态的，而在batch GD中通常用0填补，而batch norm在这种行为中会收到干扰。</div><div class="notion-text notion-block-5623936163f482229570818ee1ff5143">2.大模型的一个batch样本量通常是非常小的，batch norm的稳定性不够</div><div class="notion-text notion-block-9dc3936163f483c8b1a6014705416bd2">而到了微调阶段，则使用以下任务规划：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-f1d3936163f48326960801cda859b967"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A254edd1f-9637-4409-8c79-c92462a361ad%3Aimage.png?table=block&amp;id=f1d39361-63f4-8326-9608-01cda859b967&amp;t=f1d39361-63f4-8326-9608-01cda859b967" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-3073936163f4837487a781ff0ab7ede3">即分类任务，蕴含任务（前提和假设），相似性判断，多选择，这些任务被编排成相同的形式。</div><div class="notion-text notion-block-bfd3936163f48389bdf7816b633768e9">我们曾在CNN的迁移学习提到过，微调的基本机制是信息高斯先验，在大模型这里，预训练阶段学习到的是语言模式中普遍的信息，得到的平滑解具有极高的泛化性，当我们进行微调时，我们不是觉得这个任务的模式和预训练的模式是接近的，而是我们希望完成这个任务的模式应当仅仅通过学习残差来完成，从而继承泛化性，同时具有解决特殊问题的能力。</div><div class="notion-text notion-block-3c23936163f48386a5e5813f9d6b1645">但是这样的结果是，由于一定程度上偏离了平滑解，因此造成相当严重的“遗忘”问题，这意味着一个通才变成了一个专家。而GPT 1通过让模型同时训练无监督来抵抗这种遗忘：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-c5c3936163f482c79f23818beb392508">这里<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>是监督任务损失函数，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>是无监督任务损失函数。并且消融实验表明，同时进行无监督学习是很重要的：</div><blockquote class="notion-quote notion-block-d3e3936163f4838aa6ce0101b97f73c6"><div>We observe that the lack of pre-training hurts performance across all the tasks, resulting in a 14.8% decrease compared to our full model.</div></blockquote><div class="notion-text notion-block-fc93936163f483989adb81f5f008c0f1">预训练微调机制的关键问题在于，尽管我们已经预训练出了高效编码的语言模型，为了解决每一个特别的问题，我们需要单独微调一个模型出来，这是十分低效的</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-65c3936163f48388b47a8198fb1094b9" data-id="65c3936163f48388b47a8198fb1094b9"><span><div id="65c3936163f48388b47a8198fb1094b9" class="notion-header-anchor"></div><a class="notion-hash-link" href="#65c3936163f48388b47a8198fb1094b9" title="GPT-2 与预训练 Prompt"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">GPT-2 与预训练 Prompt</span></span></h4><div class="notion-text notion-block-9b63936163f483cbbd6101d24d4ce1c5">为了更进一步解决GPT 1的预训练-微调范式带来的低效率问题，以及更进一步缓解由微调造成的遗忘问题，OpenAI在2019年提出了GPT 2：</div><div class="notion-row"><a class="notion-bookmark notion-block-ef63936163f4837cbbbf019b63f4b534" href="https://www.semanticscholar.org/paper/Language-Models-are-Unsupervised-Multitask-Learners-Radford-Wu/9405cc0d6169988371b2755e573cc28650d14dfe" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">[PDF] Language Models are Unsupervised Multitask Learners | Semantic Scholar</div><div class="notion-bookmark-description">It is demonstrated that language models begin to learn these tasks without any explicit supervision when trained on a new dataset of millions of webpages called WebText, suggesting a promising path towards building language processing systems which learn to perform tasks from their naturally occurring demonstrations. Natural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets. We demonstrate that language models begin to learn these tasks without any explicit supervision when trained on a new dataset of millions of webpages called WebText. When conditioned on a document plus questions, the answers generated by the language model reach 55 F1 on the CoQA dataset matching or exceeding the performance of 3 out of 4 baseline systems without using the 127,000+ training examples. The capacity of the language model is essential to the success of zero-shot task transfer and increasing it improves performance in a log-linear fashion across tasks. Our largest model, GPT-2, is a 1.5B parameter Transformer that achieves state of the art results on 7 out of 8 tested language modeling datasets in a zero-shot setting but still underfits WebText. Samples from the model reflect these improvements and contain coherent paragraphs of text. These findings suggest a promising path towards building language processing systems which learn to perform tasks from their naturally occurring demonstrations.</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Fcdn.semanticscholar.org%2Fd5a7fc2a8d2c90b9%2Fimg%2Ffavicon-196x196.png?table=block&amp;id=ef639361-63f4-837c-bbbf-019b63f4b534&amp;t=ef639361-63f4-837c-bbbf-019b63f4b534" alt="[PDF] Language Models are Unsupervised Multitask Learners | Semantic Scholar" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://www.semanticscholar.org/paper/Language-Models-are-Unsupervised-Multitask-Learners-Radford-Wu/9405cc0d6169988371b2755e573cc28650d14dfe</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Fwww.semanticscholar.org%2Fimg%2Fsemantic_scholar_og.png?table=block&amp;id=ef639361-63f4-837c-bbbf-019b63f4b534&amp;t=ef639361-63f4-837c-bbbf-019b63f4b534" alt="[PDF] Language Models are Unsupervised Multitask Learners | Semantic Scholar" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-59e3936163f483b0998281b705746b80">GPT 2使用的模型为改进版的Transformer，将Layer Normalization提到shortcut之前，改进shortcut的初始化。</div><div class="notion-text notion-block-a753936163f4839abdec016bc10ede5d">GPT2的关键贡献是直接抛弃原有的预训练微调架构，提出预训练-提示架构。这是因为研究人员发现，当数据规模到达一定程度的时候，模型能够学会那些语言结构中相当高级的规律和语义行为。因此，我们只需要让模型在推理过程中自回归任务模式即可。</div><div class="notion-text notion-block-7f23936163f483418eca0156eedbecfd">GPT-2 的关键变化是进一步弱化“为每个任务单独设计监督微调格式”的思路，转而强调 <b>few-shot prompting</b>：许多 NLP 任务都可以被改写成语言模型续写问题。也就是说，只要把任务描述、输入和期望输出组织成一段文本，模型就可以通过继续生成文本来完成任务。</div><div class="notion-text notion-block-4003936163f483c6bf21014de21f587c">这就是 prompt 的基本思想：不一定改变模型结构，也不一定为每个任务重新训练一个分类头，而是通过自然语言提示，把任务变成模型已经熟悉的“根据上下文预测后续文本”。</div><div class="notion-text notion-block-21d3936163f4821981268176fdec59a0">第一个 few-shot 例子是情感分类：</div><div class="notion-text notion-block-afc3936163f48226b3010164ac89afae">模型如果继续生成：</div><div class="notion-text notion-block-1a03936163f4829295dc814dbc2b93e4">就相当于完成了情感分类。</div><div class="notion-text notion-block-da23936163f48319ac0a81b7ef333260">第二个 few-shot 例子是翻译。它的形式不是只写一句“请翻译”，而是先给出少量示范，例如：</div><div class="notion-text notion-block-15a3936163f483cc95cb01162f51c21a">这就是说，我们重复几次相同的模式，例如中文1-英文1，中文2-英文2，中文3-英文3，当我们再输入中文4时，模型将会找到前三个对应当中的模式结构，给出英文4，我们称这种方式为提示(Prompt)。</div><div class="notion-text notion-block-f7f3936163f483be9207810d64aecdd5">这两个例子说明，few-shot prompting 的核心不是训练新参数，而是把少量示范直接放进上下文，让模型通过上下文学习任务格式。Prompt 的本质不是一个额外的算法，而是一种任务表达方式：把原来需要专门建模的任务，转写成语言模型可以直接续写的文本格式。GPT-2 的重要意义就在于，它展示了大规模预训练语言模型可以通过 prompt 在许多任务上表现出一定的 zero-shot 能力。</div><div class="notion-text notion-block-fa33936163f4836c8ebb81928742a8b4">这种思路看似很简单，但是在那个时代，人们对语言模型的理解还在外推语法结构的语义连续性上，对于非常高级的逻辑和推理是几乎没有理解的。并且，通过更新权重来学习任务能力是主导思路。通过自回归Prompt来实现多任务是一次根本的范式转换，但是不难理解，相比于微调模型，GPT 2(1.5B参数量)的精度是远远不够的。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-4673936163f48335ab1c813bbe0bad5f" data-id="4673936163f48335ab1c813bbe0bad5f"><span><div id="4673936163f48335ab1c813bbe0bad5f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4673936163f48335ab1c813bbe0bad5f" title="GPT 3 and Scaling Law"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">GPT 3 and Scaling Law</span></span></h4><div class="notion-text notion-block-1893936163f482deab0f0178bd808aae">一种直接的想法就是，扩大模型的规模和训练集的规模，看看模型是否在这个过程中任务的准确率是否会有提升，这就是OpenAI在2020年提出GPT 3的思路：</div><div class="notion-row"><a class="notion-bookmark notion-block-ad53936163f482bcb64381fe76e5eeaf" href="https://arxiv.org/abs/2005.14165" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Language Models are Few-Shot Learners</div><div class="notion-bookmark-description">Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=ad539361-63f4-82bc-b643-81fe76e5eeaf&amp;t=ad539361-63f4-82bc-b643-81fe76e5eeaf" alt="Language Models are Few-Shot Learners" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2005.14165</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=ad539361-63f4-82bc-b643-81fe76e5eeaf&amp;t=ad539361-63f4-82bc-b643-81fe76e5eeaf" alt="Language Models are Few-Shot Learners" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-0473936163f482dea52a8157a3f45b8f">模型上它在GPT2的基础上引入稀疏注意力以减少计算量。通过训练8个模型观察scaling规律，它直接给出了这样一个实验数据：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-4a63936163f482c8b20601195d164012"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Aa9efe38b-1ea9-449f-9183-f2189a1917ee%3Aimage.png?table=block&amp;id=4a639361-63f4-82c8-b206-01195d164012&amp;t=4a639361-63f4-82c8-b206-01195d164012" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-e603936163f4823f829f0145f859412b">可以看到，对于1.3B Params的模型，在对话中给予prompt带来的提升是不大的。但是对于更大的模型，就开始有稳定的提升，而对于175B这样的超大模型，你会看到自然语言prompt会增强prompt的效果，即自然语言本身的结构信息进一步传入了。</div><div class="notion-text notion-block-22b3936163f483f5ad5781c06852609d">GPT3在验证scaling law的同时也发现了无限制的scaling会带来的严重问题：</div><ul class="notion-list notion-list-disc notion-block-ea23936163f4821d86498154f05f0b36"><li><b>社会偏见与毒性</b>:模型不仅会统计到语言的高级规律，也会统计到社会偏见，会被数据集中的错误信息污染</li></ul><ul class="notion-list notion-list-disc notion-block-d2c3936163f483638e2901eb7dcf040b"><li><b>幻觉</b>：模型在外推时会创造一些统计相似的但是虚假的信息</li></ul><ul class="notion-list notion-list-disc notion-block-6583936163f483e68f8281d470ffbd1a"><li>仍然缺乏推理，逻辑等更加抽象和根本的功能，这说明了纯粹文本统计结构的有限性和边际效益递减</li></ul><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-0b73936163f48215bd5b81a2611ab4f4" data-id="0b73936163f48215bd5b81a2611ab4f4"><span><div id="0b73936163f48215bd5b81a2611ab4f4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0b73936163f48215bd5b81a2611ab4f4" title="掩码语言建模与 BERT"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">掩码语言建模与 BERT</span></span></h3><div class="notion-text notion-block-2c23936163f4829699ff014fe0ad8c46">掩码语言建模将语言建模为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-4313936163f4828abea30100f0a8759b">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 token <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 被遮挡，模型需要根据其余上下文恢复它。这个思路最成功的实现之一就是 BERT。BERT 只使用 Transformer 的编码器，因此它本质上是一个编码模型，而不是像 GPT 那样的生成模型。</div><div class="notion-row"><a class="notion-bookmark notion-block-1163936163f4832baa2e81b078fe4789" href="https://arxiv.org/abs/1810.04805" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">BERT: Pre-training of Deep Bidirectional Transformers for Language...</div><div class="notion-bookmark-description">We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=11639361-63f4-832b-aa2e-81b078fe4789&amp;t=11639361-63f4-832b-aa2e-81b078fe4789" alt="BERT: Pre-training of Deep Bidirectional Transformers for Language..." loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/1810.04805</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=11639361-63f4-832b-aa2e-81b078fe4789&amp;t=11639361-63f4-832b-aa2e-81b078fe4789" alt="BERT: Pre-training of Deep Bidirectional Transformers for Language..." loading="lazy" decoding="async"/></div></a></div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-fda3936163f4828b9307818c07a6ca56" data-id="fda3936163f4828b9307818c07a6ca56"><span><div id="fda3936163f4828b9307818c07a6ca56" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fda3936163f4828b9307818c07a6ca56" title="掩码重建"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">掩码重建</span></span></h4><div class="notion-text notion-block-b4b3936163f483f0b4fb01f9fd0ce7ae">BERT 被训练为一个编码器，它的直接输出不是下一个词，而是一组上下文特征表示。</div><div class="notion-text notion-block-1593936163f482daa0a30130c9d86170">在掩码重建任务中，模型需要把被遮挡位置还原成词表中的某个 token。因此 BERT 会在编码器输出之后接一个线性层和 softmax 层，把隐藏表示投射回词表空间。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-e423936163f48303b837015cbf3f96fb"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:384px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Aa7e0e2db-e2e3-494f-a9ad-51685753e4ce%3Aimage.png?table=block&amp;id=e4239361-63f4-8303-b837-015cbf3f96fb&amp;t=e4239361-63f4-8303-b837-015cbf3f96fb" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-58e3936163f483b4b91301b15614f4b4">除此之外，BERT 还通过 NSP（Next Sentence Prediction，下一句预测）任务训练编码器，让编码器形成更强的全局语义视野。NSP 的形式是给出两个句子 A 和 B，让 BERT 判断 B 是否真的是 A 的下一句。</div><div class="notion-text notion-block-be63936163f4820fad8781e9246e3a1e">当 BERT 被微调用于下游任务时，通常保留编码器主体，只把输出头换成具体任务需要的形式。</div><div class="notion-text notion-block-7213936163f48361b0bb81d79fff3680">把预训练、任务头替换和下游微调三个过程合在一起，就得到如下流程：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-fa73936163f4821fa708018e9c4a2918"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A554a389c-a21d-4400-851e-36cad53164f6%3Aimage.png?table=block&amp;id=fa739361-63f4-821f-a708-018e9c4a2918&amp;t=fa739361-63f4-821f-a708-018e9c4a2918" alt="notion image" loading="lazy" decoding="async"/></div></figure><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-2c23936163f4835ebef10159a76410b2" data-id="2c23936163f4835ebef10159a76410b2"><span><div id="2c23936163f4835ebef10159a76410b2" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2c23936163f4835ebef10159a76410b2" title="Segment Embeddings"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Segment Embeddings</span></span></h4><div class="notion-text notion-block-dfc3936163f4821a91bc81c7fc8d35b7">为了同时支持掩码重建和下一句预测两个任务，BERT 扩展了 Embedding 的含义。</div><div class="notion-text notion-block-b013936163f48211a885018a976b698f">除了用于提供词语语义空间的 Token Embedding，以及告诉模型序列位置的 Position Embedding，BERT 还引入 Segment Embeddings，用来告诉模型：当前输入中包含两个不同的句子片段。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-aeb3936163f483b986eb01a2884b0af5"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Aaf60b6e1-b0b6-400b-80cf-6cf722a80cff%3Aimage.png?table=block&amp;id=aeb39361-63f4-83b9-86eb-01a2884b0af5&amp;t=aeb39361-63f4-83b9-86eb-01a2884b0af5" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-8b53936163f48221b1ae8170a5061a8d">在微调阶段，Segment Embeddings 也可以用来区分不同输入片段，例如在问答任务中区分问题和段落。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-9a03936163f482c4bb9081d740e84afc" data-id="9a03936163f482c4bb9081d740e84afc"><span><div id="9a03936163f482c4bb9081d740e84afc" class="notion-header-anchor"></div><a class="notion-hash-link" href="#9a03936163f482c4bb9081d740e84afc" title="分类 token"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">分类 token</span></span></h4><div class="notion-text notion-block-3903936163f482ed858a014192df206d">许多任务，包括 NSP，并不需要输出一个完整序列，而是只需要输出一个概率或一个类别。这意味着模型需要把整段序列的特征汇总成一个全局表示。</div><div class="notion-text notion-block-e3b3936163f4835ba9e1815c72e0ff26">一个直接想法是用全连接层汇总所有 token 的 embedding，但 BERT 使用的是 CLS token。CLS 可以理解为一种通过自注意力学习出来的隐式序列摘要。</div><div class="notion-text notion-block-5353936163f4824da47b8153b6d8782e">CLS 是每个输入序列开头的特殊 token，用来承载全局信息。通过自注意力机制，所有 token 都会与 CLS 发生交互。因此，如果任务要求 CLS 具有全局视野，例如 NSP，训练过程会推动 CLS 学会汇总整段输入的信息。</div><div class="notion-text notion-block-d363936163f482148205019fc9ee38c2">同时，其他 token 也可以从 CLS 中读取全局信息，这在一定程度上有助于表示学习。</div><div class="notion-text notion-block-2903936163f483d18fde81d9c6d8bd5e"><b>不过，后续研究表明，CLS token 本身并不一定带来很大的性能提升，它更像是一种简洁的全局读出设计。</b></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-a5c3936163f483d0b1c0812800c351f6" data-id="a5c3936163f483d0b1c0812800c351f6"><span><div id="a5c3936163f483d0b1c0812800c351f6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#a5c3936163f483d0b1c0812800c351f6" title="Transformer In CV"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Transformer In CV</span></span></h2><div class="notion-text notion-block-5973936163f482838c610151055e2c84">Transformer在MLP领域取得一定成功后，就有人开始将其迁移到计算机视觉领域，形成了CONV和Self Attention的混合架构，也有人完全抛弃了CONV，但是它们由于使用了稀疏注意力，导致计算效率低下，因此ResNet在大部分时间内仍然是SOTA。为什么NLP的成功如此难以迁移到图像，我们从自回归图像的发展过程就可以看出来。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-e9b3936163f4837e8e7f81ca2ff3ad84" data-id="e9b3936163f4837e8e7f81ca2ff3ad84"><span><div id="e9b3936163f4837e8e7f81ca2ff3ad84" class="notion-header-anchor"></div><a class="notion-hash-link" href="#e9b3936163f4837e8e7f81ca2ff3ad84" title="AutoRegressive Image"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">AutoRegressive Image</span></span></h3><div class="notion-text notion-block-3d73936163f48374a9900131538e0c6d">自回归图像早有历史，早在2017年就Google就尝试过使用RNN，并且像素级别回归图像，但是受限于RNN架构本身的限制，未能取得理想成果。</div><div class="notion-text notion-block-e3c3936163f482b5a2fa81ef6e7939a6">2018年，Image Transformer用Transformer来进行回归，但是计算量过大，同样有限。</div><div class="notion-row"><a class="notion-bookmark notion-block-eb93936163f4828d83ca81a3521ea07a" href="https://arxiv.org/abs/1802.05751" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Image Transformer</div><div class="notion-bookmark-description">Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=eb939361-63f4-828d-83ca-81a3521ea07a&amp;t=eb939361-63f4-828d-83ca-81a3521ea07a" alt="Image Transformer" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/1802.05751</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=eb939361-63f4-828d-83ca-81a3521ea07a&amp;t=eb939361-63f4-828d-83ca-81a3521ea07a" alt="Image Transformer" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-7d23936163f48331a979815489cf0333">Open AI将图像下采样，在采样后的图像上自回归以缓解困难，提出了iGPT，但是牺牲了图像质量</div><div class="notion-row"><a class="notion-bookmark notion-block-66d3936163f483c7a8ba81ca12b29132" href="https://www.semanticscholar.org/paper/Generative-Pretraining-From-Pixels-Chen-Radford/bc022dbb37b1bbf3905a7404d19c03ccbf6b81a8" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">[PDF] Generative Pretraining From Pixels | Semantic Scholar</div><div class="notion-bookmark-description">It is found that a GPT-2 scale model learns strong image representations as measured by linear probing, ﬁne-tuning, and low-data classiﬁcation, despite training on low-resolution ImageNet without labels. Inspired by progress in unsupervised representation learning for natural language, we examine whether similar models can learn useful representations for images. We train a sequence Transformer to auto-regressively predict pixels, without incorporating knowledge of the 2D input structure. Despite training on low-resolution ImageNet without labels, we ﬁnd that a GPT-2 scale model learns strong image representations as measured by linear probing, ﬁne-tuning, and low-data classiﬁcation. On CIFAR-10, we achieve 96.3% accuracy with a linear probe, outperforming a supervised Wide ResNet, and 99.0% accuracy with full ﬁne-tuning, matching the top supervised pre-trained models. An even larger model trained on a mix-ture of ImageNet and web images is competitive with self-supervised benchmarks on ImageNet, achieving 72.0% top-1 accuracy on a linear probe of our features.</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Fcdn.semanticscholar.org%2Fd5a7fc2a8d2c90b9%2Fimg%2Ffavicon-196x196.png?table=block&amp;id=66d39361-63f4-83c7-a8ba-81ca12b29132&amp;t=66d39361-63f4-83c7-a8ba-81ca12b29132" alt="[PDF] Generative Pretraining From Pixels | Semantic Scholar" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://www.semanticscholar.org/paper/Generative-Pretraining-From-Pixels-Chen-Radford/bc022dbb37b1bbf3905a7404d19c03ccbf6b81a8</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Fwww.semanticscholar.org%2Fimg%2Fsemantic_scholar_og.png?table=block&amp;id=66d39361-63f4-83c7-a8ba-81ca12b29132&amp;t=66d39361-63f4-83c7-a8ba-81ca12b29132" alt="[PDF] Generative Pretraining From Pixels | Semantic Scholar" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-05a3936163f4833b85c78125a9bc8eae">这一切的问题都在于，对于同等语义量，图像的tokens长度远远长于文本，自回归虽然在自然语言成功了，在图像方面则不是这样，仅仅使用Tranformer解码器是根本不够的。</div><div class="notion-text notion-block-3463936163f48390b67501d49d94d5e2">并且，图像结构并不保证具有良好的顺序条件概率结构，这种建模在本质上是一种扭曲的近似。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-7ce3936163f48269a9b501577315cafe" data-id="7ce3936163f48269a9b501577315cafe"><span><div id="7ce3936163f48269a9b501577315cafe" class="notion-header-anchor"></div><a class="notion-hash-link" href="#7ce3936163f48269a9b501577315cafe" title="视觉 Transformer"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">视觉 Transformer</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-26a3936163f483e6a21701db51696f3d" data-id="26a3936163f483e6a21701db51696f3d"><span><div id="26a3936163f483e6a21701db51696f3d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#26a3936163f483e6a21701db51696f3d" title="架构"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">架构</span></span></h4><div class="notion-text notion-block-ef23936163f482cbb478811f80a6d348">2020年Google团队提出Vision Transformer，通过引入Patch思路，使用最基本的全局注意力架构，取得了成功：</div><div class="notion-row"><a class="notion-bookmark notion-block-2513936163f483cc850401aecd2eb422" href="https://arxiv.org/abs/2010.11929" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale</div><div class="notion-bookmark-description">While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=25139361-63f4-83cc-8504-01aecd2eb422&amp;t=25139361-63f4-83cc-8504-01aecd2eb422" alt="An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2010.11929</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=25139361-63f4-83cc-8504-01aecd2eb422&amp;t=25139361-63f4-83cc-8504-01aecd2eb422" alt="An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-d4b3936163f482d0bb2b811e3889258c">基本思路是，将图像切成nxn份，对每一份展平为向量，并linearly embed，可学习PE，然后把<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>个当作tokens传入正常的transformer编码器中，如下图</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-8763936163f483f8994b01337d3f9699"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ab50fda51-3566-4186-a63a-f04a5e707aa1%3Aimage.png?table=block&amp;id=87639361-63f4-83f8-994b-01337d3f9699&amp;t=87639361-63f4-83f8-994b-01337d3f9699" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-df73936163f482c89aaa81c3ff74186e">可以看到，这是一个监督学习任务，结构中采用Pre-LayerNorm，激活函数使用GELU(Gauss Error Linear Unit)，它是现代发现的在某些问题上比RELU更好的激活函数，定义为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>为高斯分布的概率累计函数，图像为</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-6ea3936163f4835db62581befd03267f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:384px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ac76cc58f-7f0c-4b5d-ae3a-a06826906f65%3Aimage.png?table=block&amp;id=6ea39361-63f4-835d-b625-81befd03267f&amp;t=6ea39361-63f4-835d-b625-81befd03267f" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-af23936163f483619bcc014d21c19918">可以看到在0点是非奇异的。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-4e13936163f483eb9e6c813460cfdd4e" data-id="4e13936163f483eb9e6c813460cfdd4e"><span><div id="4e13936163f483eb9e6c813460cfdd4e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4e13936163f483eb9e6c813460cfdd4e" title="更高分辨率"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">更高分辨率</span></span></h4><div class="notion-text notion-block-1ae3936163f4820ba29b019a40e42243">假设模型是在 <code class="notion-inline-code">224x224</code> 分辨率、<code class="notion-inline-code">16x16</code> 的 Patch 上预训练的（序列长度 = <code class="notion-inline-code">(224/16)^2 = 196</code>）。在微调或推理时，如果输入图像是 <code class="notion-inline-code">384x384</code>，我们仍然使用 <code class="notion-inline-code">16x16</code> 的 Patch。此时，序列长度变为 <code class="notion-inline-code">(384/16)^2 = 576</code>。这比预训练时的 <code class="notion-inline-code">196</code> 要长得多。</div><div class="notion-text notion-block-e473936163f482ca8e2281e96b3fa132">Transformer 本身可以处理更长的序列（只要算力允许），但问题在于<b>位置嵌入</b>。在预训练中，模型为序列中的 <b>前 196 个位置</b> 学习了特定的位置编码。现在序列长度变成了 576，<b>第 197 到第 576 个位置</b> 没有对应的、经过训练的位置编码。</div><div class="notion-text notion-block-abb3936163f4820a9181010b0cab35d7">ViF的思路是，位置编码应该对应于图像中的<b>空间位置</b>，而不是序列中的索引。将预训练的位置嵌入矩阵视为一个<b>低分辨率的 2D 网格，</b>然后，使用 <b>2D 插值算法</b>（如双线性插值）将这个 <code class="notion-inline-code">14x14</code> 的网格<b>上采样</b>到目标分辨率所需的网格大小（例如 <code class="notion-inline-code">24x24</code>，因为 <code class="notion-inline-code">384/16=24</code>）。
这样，每个 Patch 的位置编码都根据其在原始图像 2D 空间中的<b>实际位置</b>进行了重新计算，保持了空间关系的连续性。</div><div class="notion-text notion-block-c983936163f4837896b181a508e62ae2">这就是ViF最关键的图像结果先验的嵌入。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-5b73936163f482189768815075bdfe39" data-id="5b73936163f482189768815075bdfe39"><span><div id="5b73936163f482189768815075bdfe39" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5b73936163f482189768815075bdfe39" title="Scaling"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Scaling</span></span></h4><div class="notion-text notion-block-6673936163f4839e8c6481d7a2181c12">想要证明ViF的成功性，必须表明，尽管数据量少时可能无法击败CNN的SOTA，但是数据量大时，可以具有非常优良的增长。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3473936163f48373892701865511a115"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A4c895c62-8837-43c5-a959-86532748fdc8%3Aimage.png?table=block&amp;id=34739361-63f4-8373-8927-01865511a115&amp;t=34739361-63f4-8373-8927-01865511a115" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-7c93936163f4827f8f5e81f358fc162d">并且Vision Transformers appear not to saturate within the range tried, motivating future scaling efforts.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-b8f3936163f4824191b781086211d878" data-id="b8f3936163f4824191b781086211d878"><span><div id="b8f3936163f4824191b781086211d878" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b8f3936163f4824191b781086211d878" title="学会成为更好的 CNN"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">学会成为更好的 CNN</span></span></h4><div class="notion-text notion-block-5673936163f483918265016a8ff9af0f">首先对于底层的Attention的PCA分析表明，它们在关注基础视觉元素</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-1073936163f483f0a5bd01c1380147e5"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:528px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ab2e65a1b-e809-4d07-8b4a-21a4ee8906ac%3Aimage.png?table=block&amp;id=10739361-63f4-83f0-a5bd-01c1380147e5&amp;t=10739361-63f4-83f0-a5bd-01c1380147e5" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-0d33936163f48391ac710123fdba89f5">并且在embedding矩阵的相似度分析中发现，Transformer仅仅通过PE学到了行列结构</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-d0b3936163f4822aa62881fe0fcd4ef8"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A57e3a58c-21a2-44aa-a705-1654c6702373%3Aimage.png?table=block&amp;id=d0b39361-63f4-822a-a628-81fe0fcd4ef8&amp;t=d0b39361-63f4-822a-a628-81fe0fcd4ef8" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-d383936163f4825c9bac014023df39f2">这就是为什么hybrid效果一般甚至更差，因为CNN的结构先验其实被网络自己学到了。然后通过注意力距离分析发现，即使在早期层，注意力也已经非常长程化了，而不需要CNN的下采样的感受野扩大过程。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-9c13936163f483d1a916013645754772"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A4a2f6ee4-df9c-44bb-8dc0-dd279796e99e%3Aimage.png?table=block&amp;id=9c139361-63f4-83d1-a916-013645754772&amp;t=9c139361-63f4-83d1-a916-013645754772" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-c353936163f48220bf5b016f04b55e5e">以上结果都表明，ViF再次验证了弱架构先验的中心思路。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-36e3936163f483478e1c81421a88b146" data-id="36e3936163f483478e1c81421a88b146"><span><div id="36e3936163f483478e1c81421a88b146" class="notion-header-anchor"></div><a class="notion-hash-link" href="#36e3936163f483478e1c81421a88b146" title="掩码自编码器"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">掩码自编码器</span></span></h3><div class="notion-row"><a class="notion-bookmark notion-block-b153936163f48303834f81fc9a29701d" href="https://arxiv.org/abs/2111.06377" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Masked Autoencoders Are Scalable Vision Learners</div><div class="notion-bookmark-description">This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=b1539361-63f4-8303-834f-81fc9a29701d&amp;t=b1539361-63f4-8303-834f-81fc9a29701d" alt="Masked Autoencoders Are Scalable Vision Learners" loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/2111.06377</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=b1539361-63f4-8303-834f-81fc9a29701d&amp;t=b1539361-63f4-8303-834f-81fc9a29701d" alt="Masked Autoencoders Are Scalable Vision Learners" loading="lazy" decoding="async"/></div></a></div><div class="notion-text notion-block-d5d3936163f482baa2c7812a33260166">掩码自编码器并不是稀奇的技术，在Bert那里就已经这样做了，通过遮挡一部分文字，然后让编码器编码未被遮挡的文字，用Linear+Softmax解码全部的文字。MAE的创新在于他解决了图像和语言两种信息的差异性所导致的问题。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-4593936163f483bd8e3581c6bb6ff47d" data-id="4593936163f483bd8e3581c6bb6ff47d"><span><div id="4593936163f483bd8e3581c6bb6ff47d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4593936163f483bd8e3581c6bb6ff47d" title="高比例掩码"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">高比例掩码</span></span></h4><ul class="notion-list notion-list-disc notion-block-67d3936163f4836bb42a810d632f457b"><li>Information density is different between language and vision</li></ul><div class="notion-text notion-block-1c73936163f483d9857001b40dfd6f17">文字是高精度编码的，图像数据是高度冗余的，这使得a missing patch can be recovered from neighboring patches with little high-level understanding of parts, objects, and scenes。</div><div class="notion-text notion-block-7c53936163f482299bd1013a97c0b21c">He Kaiming的解决方法是掩码大部分区域：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-1b33936163f4821e8e3501652d7dea7d"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:528px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ab58e85ab-922f-4d42-b822-976d789e2650%3Aimage.png?table=block&amp;id=1b339361-63f4-821e-8e35-01652d7dea7d&amp;t=1b339361-63f4-821e-8e35-01652d7dea7d" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-c2a3936163f48261b2d181c0295dc278">即让模型需要花费更多的功夫恢复图像，从而建立更高级的表征。这造成一个双赢局面，高度掩码减小了内存占用，并且同时提高了训练效果。</div><div class="notion-text notion-block-ea53936163f482b19d4a013479840326">并且得到</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-f5a3936163f483dc9d45811f7985daf2"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:528px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ab7b10753-472f-4a36-a977-45eff1763ad4%3Aimage.png?table=block&amp;id=f5a39361-63f4-83dc-9d45-811f7985daf2&amp;t=f5a39361-63f4-83dc-9d45-811f7985daf2" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-9393936163f482c5867201093d1feae2">这就是为什么He Kaiming说：</div><blockquote class="notion-quote notion-block-bf53936163f48347bfa5811b029450d2"><div>The model infers missing patches to produce different, yet plausible, outputs (Figure 4). It makes sense of the gestalt of objects and scenes, which cannot be simply completed by extending lines or textures. We hypothesize that this reasoning-like behavior is linked to the learning of useful representations.</div></blockquote><div class="notion-text notion-block-86d3936163f482b39d7a019ea5232680">同样文章还研究了不同掩码方法的影响</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-eae3936163f4820dbd970194438cd09d"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:576px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A9b770f10-b04d-4412-9b4e-7f6ee1aebe00%3Aimage.png?table=block&amp;id=eae39361-63f4-820d-bd97-0194438cd09d&amp;t=eae39361-63f4-820d-bd97-0194438cd09d" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-d663936163f483a7b50601e5bd70d692">很容易理解，随机掩码的效果最好，模型不会因为学习到掩码规律而偷懒。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-6d03936163f483e58e0b013b15f64fc5" data-id="6d03936163f483e58e0b013b15f64fc5"><span><div id="6d03936163f483e58e0b013b15f64fc5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6d03936163f483e58e0b013b15f64fc5" title="轻量级解码器"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">轻量级解码器</span></span></h4><ul class="notion-list notion-list-disc notion-block-d3b3936163f48294b891815f9140957a"><li>The autoencoder’s decoder, which maps the latent representation back to the input, plays a different role between reconstructing text and images</li></ul><div class="notion-text notion-block-9763936163f482849a7301f03fd7e9b1">这指的是bert仅仅通过线性层+softmax就可以讲特征投射回语言空间，因为文字是离散的，这是个classification问题，而对于图像，需要引入完整的transformer decoder，但是必须是不对称的，our decoder is lightweight and reconstructs the input from the latent representation along with mask tokens如下图</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-7503936163f483508ca80168e7665841"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Af96dafb8-2357-41ea-a1d0-ca5a655649bc%3Aimage.png?table=block&amp;id=75039361-63f4-8350-8ca8-0168e7665841&amp;t=75039361-63f4-8350-8ca8-0168e7665841" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-8773936163f4824f9f4c01083cb78722">效果是显著的 With a vanilla ViT-Huge model, we achieve 87.8% accuracy when finetuned on ImageNet-1K. This outperforms all previous results that use only ImageNet-1K data.</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-2b13936163f483cb9c1281b88c57847c" data-id="2b13936163f483cb9c1281b88c57847c"><span><div id="2b13936163f483cb9c1281b88c57847c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2b13936163f483cb9c1281b88c57847c" title="Scaling 扩展"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Scaling 扩展</span></span></h4><div class="notion-text notion-block-a383936163f48315ad5a81abf138eae0">最后是Scaling方面</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-3513936163f48217b21601db0398626d"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A78b71fe6-3d02-454a-98d0-484762118f7c%3Aimage.png?table=block&amp;id=35139361-63f4-8217-b216-01db0398626d&amp;t=35139361-63f4-8217-b216-01db0398626d" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-8a33936163f482a09fd2014de78d116f">性能仍然在增长，这表明了强大的scaling前景。</div><div class="notion-blank notion-block-0463936163f48223ab60816d1c7e8102"> </div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Phase Transition Theory]]></title>
            <link>https://blog.xiangsiqi.site/PhaseTransition</link>
            <guid>https://blog.xiangsiqi.site/PhaseTransition</guid>
            <pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39d3936163f48087a5ffca135cddb69d"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-default notion-h-indent-0 notion-block-39d3936163f4807ab132eaaa0c4fa878" data-id="39d3936163f4807ab132eaaa0c4fa878"><span><div id="39d3936163f4807ab132eaaa0c4fa878" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4807ab132eaaa0c4fa878" title="Ising Model and Mean Field Approximation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Ising Model and Mean Field Approximation</span></span></h2><div class="notion-text notion-block-39d3936163f481f9b92ff3ba6ab5f937">伊辛模型是相变理论中最小而最有力的模型之一。它只保留每个格点上二值自旋的自由度，却已经能够展示有序相、无序相、自发对称性破缺、临界点、临界指数和涨落的重要性。后面的 Landau 理论与 Landau-Ginzburg 理论，很多概念都可以先在这个模型里看清楚。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f48115b87cdee8a20248f5" data-id="39d3936163f48115b87cdee8a20248f5"><span><div id="39d3936163f48115b87cdee8a20248f5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f48115b87cdee8a20248f5" title="1. 模型定义"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. 模型定义</span></span></h3><div class="notion-text notion-block-39d3936163f48121a409fa0efb11ad8b">在一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 维晶格上，每个格点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 放置一个自旋变量</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481acb252e2386c5a1f69">最近邻铁磁 Ising 模型的哈密顿量为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-b8130fcf1b3041d39ebe177cf62cedac">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示最近邻格点对，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是相邻自旋之间的耦合常数，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是外磁场。<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，同向排列的相邻自旋能量更低，因此低温下系统倾向于形成铁磁有序。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f4819ca906f83bf35e869f" data-id="39d3936163f4819ca906f83bf35e869f"><span><div id="39d3936163f4819ca906f83bf35e869f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4819ca906f83bf35e869f" title="2. 配分函数与热平均"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. 配分函数与热平均</span></span></h3><div class="notion-text notion-block-39d3936163f4816298aec023a5e952bf">对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个格点的 Ising 系统，正则配分函数为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48183886bc5ca5c1f5fb8">把哈密顿量代入，可写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f480a79888f79e57329ad7">也写为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48194b2c0de60fee13d8b">自由能为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48124aed4e4ff966d4f32">磁化强度，也就是 Ising 模型的序参量，定义为单位格点平均自旋：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-7fc718a221084cc6b4064f0ec1aa41f6">由于外场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 正是耦合到总磁矩 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的源项，所以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可以直接由配分函数对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 求导得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-88cf4f6aaec444b19300e1a035ca7ebf">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-4b8127af21b84afb974d21c64a023cbf">这就是“序参量和配分函数的关系”。磁化率也同样来自二阶导数：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-57a39c6c93c64571b24331709cc715e8">所以磁化率本质上是总磁矩涨落。临界点附近 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 发散，等价于大尺度磁化涨落变得异常强。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-3a33936163f48166b3c0c5d7d794240c" data-id="3a33936163f48166b3c0c5d7d794240c"><span><div id="3a33936163f48166b3c0c5d7d794240c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3a33936163f48166b3c0c5d7d794240c" title="3. 自由能的凸性"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 自由能的凸性</span></span></h3><div class="notion-text notion-block-3a33936163f481c895a9dd80444354e3">这一节关心的不是某个具体近似，而是配分函数本身带来的热力学稳定性。自由能的凸性告诉我们：哪些响应函数必须非负，哪些地方可能出现真正的相变。</div><div class="notion-text notion-block-3a33936163f4816ba4efe53ee1203a4c">先给定义。设 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是定义在区间上的函数。如果对任意 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，都有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f48172b8b3fa90c3ef8ef0">则称 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是凸函数。如果不等号方向反过来，则称为凹函数。若二阶导数存在，凸性等价于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，凹性等价于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-3a33936163f4810499f4e647913972b7">为了证明配分函数的凸性，把参数统一写成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。若哈密顿量中某个观测量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 线性耦合，则有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481a8a45dd36ee8ada3c7">例如对 Ising 模型的外场变量，可以取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-3a33936163f481a39899c525bcd57c48">取两个参数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481edb35dd8876978a741">对上式使用 Holder 不等式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f4817085cdf51a4411b7d1">于是得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f48105a079f16b7930721c">两边取对数，便有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481c1a51fce924a05246b">因此 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对线性耦合参数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是凸函数。若二阶导数存在，这等价于</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481a7b6d7c53d8c843674">这条公式的物理含义非常直接：凸性来自涨落的非负性。方差不可能为负，所以相应的响应函数也不可能为负。</div><div class="notion-text notion-block-3a33936163f4812dbbd3fbdba9f0686f">对外场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f4812faa9fd16a1fdaf41b">由于</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481028fecd904b18a45d4">所以自由能密度对外场是凹的：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481759254f43e282274bc">磁化强度满足 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此磁化率为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f48149bd27d38f9dcfe92d">对温度也有类似稳定性。熵密度定义为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f4813aa416c362a321e97b">定外场比热密度为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481e19c07d8fdbeace699">在正则系综中，这也等价于能量涨落：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a33936163f481f0b480fd0d9492faa1">因此 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。换句话说，自由能对温度是凹的，比热非负；自由能对外场按上面的符号约定也是凹的，磁化率非负。</div><div class="notion-text notion-block-3a33936163f48130bac0cd8c3caada13">这和相变的关系在于：有限系统的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是有限项指数函数之和，所以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 通常是光滑的；只有热力学极限中，自由能才可能在保持凸性/凹性的同时出现非解析点。</div><div class="notion-text notion-block-3a33936163f481818311e0dbe0b4db21">一级相变对应自由能一阶导数跳变，例如 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 不连续。连续相变中一阶导数可以连续，但二阶导数可能发散或不连续，例如 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在临界点附近变得很大甚至发散。</div><div class="notion-text notion-block-3a33936163f4813d9f05e8c0de82a2d3">所以，自由能的凸性不是形式数学装饰，而是热力学稳定性的表达：比热和磁化率来自自由能的二阶导数，也来自相应物理量的涨落；相变则是热力学极限中这些导数失去普通解析性的地方。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-4d0af999216647c4809879e25d100e7c" data-id="4d0af999216647c4809879e25d100e7c"><span><div id="4d0af999216647c4809879e25d100e7c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4d0af999216647c4809879e25d100e7c" title="4. 有序、无序与自发对称性破缺"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. 有序、无序与自发对称性破缺</span></span></h3><div class="notion-text notion-block-0fea6349dfa242dfaf7b0ac98b414884">在零外场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，哈密顿量在全局翻转</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-086e284a03274a1d8cab655914043173">下不变。因此模型本身不偏好“全向上”或“全向下”。高温下热涨落主导，自旋取向混乱，平衡态保持这个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对称性，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-a1d061e58b7b48ec9bc6e987149b6070">低温下能量主导，相邻自旋倾向于同向排列。系统会在两个等价的有序态中选一个：<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。哈密顿量仍然对称，但具体平衡态不再对称，这就是自发对称性破缺。</div><div class="notion-text notion-block-956db1d6fca24a2ebd5ba8b744053e8d">严格地说，在有限系统且 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，由于两个有序态权重相等，配分函数给出的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 仍然为零。真正的自发磁化需要按如下顺序取极限：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-c84a07a156a54ce8acdd6d47fcf02460">先取热力学极限，再让外场趋于零，系统才会选择一个破缺对称性的纯态。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-50451dd5ee6446d4bde1d2bf460b4988" data-id="50451dd5ee6446d4bde1d2bf460b4988"><span><div id="50451dd5ee6446d4bde1d2bf460b4988" class="notion-header-anchor"></div><a class="notion-hash-link" href="#50451dd5ee6446d4bde1d2bf460b4988" title="5. 平均场理论"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. 平均场理论</span></span></h3><div class="notion-text notion-block-b1f402c213274b4281d8be8765cafe28">平均场理论的基本近似是：不再逐个追踪某个自旋周围邻居的真实取值，而是把邻居替换成平均磁化强度 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。如果每个格点有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个最近邻，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-e524d2dfecc945d19426585b512736b8">这个公式来自把每个自旋分解成“平均值 + 涨落”：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9d7dc061e5704afab7a7eb234eba7cbc">于是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3111e9ded9aa4fbd8ff8ffd3a1b80b79">展开为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-724af44bdc05492cadb5051b8b3a2f05">整理后三项得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-8729bf609c4c42e9b2f5e51315e4df84">平均场近似的关键就是忽略二阶涨落关联项：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-8895f1ed7d4e49afa44e72335879809f">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-b3b6250d5a3048bc82790854e5c65d63">也就是说，平均场理论并不是说每个自旋完全等于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而是说相邻自旋之间的涨落关联 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 被忽略了，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这正是它后来在临界点附近失效的根源：临界点附近涨落关联会变得长程而强烈，不能再被当作小量丢掉。</div><div class="notion-text notion-block-7d3dfc1391904038b0cf5143d464c34c">代入哈密顿量后得到平均场哈密顿量</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-48228bc941e04fb7942f25ac01761bb1">这个式子的代入过程如下。从原哈密顿量</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-04f84b1881744f038cb086749f04bf93">出发，将</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-7aa7061d51174fdeaed2783404203c3f">代入相互作用项：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-0f773284bf3c4163ad24d5f112d6ab8c">拆开得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-af438c3a32024d9ab01a78703fa1c4b5">现在需要处理两个最近邻键求和。若每个格点有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个最近邻，则在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中，每个格点自旋 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 会被数到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 次，因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-6fdbe325e5a74cbd9999911c8728143f">最近邻键的总数是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a86db8d8744949be8995f45a5aae4512">这里先从每个格点出发数出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 条连接，但每条键连接两个格点，被重复数了两次，所以要除以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-17d6a674585b4d4bb0456ecf8bf692ee">再加上外场项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3ab46179299a4b1ca3c47d89f5a06452">所以最后一项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 来自最近邻键数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，也可以理解为对平均场相互作用能双计数的修正。于是平均场配分函数变成独立单自旋配分函数的乘积：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a35a91b6a9e649b2b0dca271d2635463">这个平均场配分函数也可以逐步看出来。由定义</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-2a5f356c21104c568917bd409042a3d1">并代入</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-6f8e162bf1b94a0da8eef5c1f02bb263">得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-b556042f5c1f4dbbbc56bbe640207ffb">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 不依赖具体自旋构型，可以提出求和号：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-ad92b61c527740a2b8b2793550b84089">再利用</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a8db24902c9342f38ce24aa01d1b2d8f">平均场哈密顿量已经没有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 这样的自旋间耦合项，所以构型求和可以分解成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个独立单自旋求和：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-eba5b6878dbf43b4984b4d561217f769">对每个格点，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-6e44ace5edfb4794a8b012c3ae8870cf">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-061764f8697643e58769596f767bf536">最终得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-10474ed40f0e4de9b5e01662da166d56">这里的核心是：常数项可以从构型求和中提出，而平均场近似使系统变成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个处在有效场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中的独立自旋。</div><div class="notion-text notion-block-86f3024813624faaaa46d4bcf15b7ab9">平均场自由能密度为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-61a8227f06f84b8c940502219f5ee0bb">由极小化条件 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，或等价地由单自旋热平均，得到自洽方程</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39d3936163f4819f96cbc0b7d8efa2cc"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://www.notion.so/image/attachment%3Abac25615-bda6-4cb5-87cf-c974e3f7e0b8%3Aising_mean_field_roots.png?table=block&amp;id=39d39361-63f4-819f-96cb-c0b7d8efa2cc&amp;t=39d39361-63f4-819f-96cb-c0b7d8efa2cc" alt="Roots of the mean-field self-consistency equation m = tanh(km), where k = βJz." loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">Roots of the mean-field self-consistency equation m = tanh(km), where k = βJz.</figcaption></div></figure><div class="notion-text notion-block-39d3936163f48113a8e4ee6b9f12cdaa">图中交点就是自洽方程 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的根，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，只有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 一个根；当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，除了 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 之外，还出现两个非零根 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这说明非零自洽解是在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 处从零解处分叉出来的。</div><div class="notion-text notion-block-39d3936163f481e3ac20e3e0c1541cc3">因此可以直接从图像理解临界点。令 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在原点附近的斜率小于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的斜率，所以两条曲线只有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 一个交点，对应顺磁解。</div><div class="notion-text notion-block-39d3936163f48107a43ac0278ed8a769">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在原点附近比 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 增长得更快，原点仍然是数学上的交点，但两侧会长出两个新的非零交点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此临界条件就是两条曲线在原点处斜率相等：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812c9fb8fc5525d3bfc2">由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481d49258db2c5f729e16">接下来计算低温侧非零根的大小。临界点附近 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 很小，对自洽方程</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48174b9e9da91ab5a26b3">做小量展开：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e8b653e292308016ea">代入得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481839b02c246f65eeaf0">移项并提取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4811cab10e45fc6722a79">除了 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 之外，非零解满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48181a62ee1303c25ba04">因为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812e87a9fdf62b1d23ca">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481478d22c5bcbae3dbd9">这给出平均场临界指数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-39d3936163f48163903af52736d18138">磁化率也可以直接从平均场自洽方程推出。这里暂时用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示外磁场，用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示配位数，也就是前文的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是单位格点磁化强度，总磁化率定义为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481beb927fe9f11262f2a">平均场自洽方程写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481d2b73cf52b40365151">对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 求导，并保持温度不变：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481428974cacb2768dd93">再用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，即 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4818fb5f1d04cab3f70be">这里最后一步取了 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在顺磁相中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，上式化为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481cba449f776a6e3bd0c">移项可得</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4810db300f5b2e6dd0f74">所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e9b313c3785e1b76c2">这里使用了 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，分母趋于零，因此磁化率发散，平均场给出临界指数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f481698d40df6c489c94db" data-id="39d3936163f481698d40df6c489c94db"><span><div id="39d3936163f481698d40df6c489c94db" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f481698d40df6c489c94db" title="6. 临界指数"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">6. 临界指数</span></span></h3><div class="notion-text notion-block-39d3936163f481f08e61fd701efcb709">前面从平均场自洽方程出发，对临界点附近的小 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 做展开，已经得到了自发磁化和磁化率的临界行为。这里的关键近似是：只在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近保留最低阶非平凡项，因此得到的是临界附近的渐近幂律，而不是全温区公式。</div><div class="notion-text notion-block-39d3936163f48154bf5ef98682e7b607">临界指数就是这些幂律中的指数。平均场理论给出的主要结果是：</div><ul class="notion-list notion-list-disc notion-block-39d3936163f4811abe49d6f66f08b8b2"><li>自发磁化：<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</li></ul><ul class="notion-list notion-list-disc notion-block-39d3936163f48131b673e4fd7372928b"><li>磁化率：<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</li></ul><ul class="notion-list notion-list-disc notion-block-39d3936163f48138b59ee867cbbbcd79"><li>临界等温线：<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</li></ul><div class="notion-text notion-block-39d3936163f481788fb0f141a06b5d50">这些指数描述的是临界点附近的奇异行为。它们比具体的临界温度更重要，因为临界指数体现了普适性：不同微观模型只要维度、对称性和相互作用范围相同，就可能共享同一组临界指数。</div><div class="notion-text notion-block-39d3936163f4817085b2e3d098b4c922">临界指数背后的物理机制，可以从磁化率、磁化涨落和关联函数的关系看得更清楚。令总磁矩为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48113947af0091a563db7">若磁化率定义为单位格点磁化强度对外场的响应，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4817c8e80c9154aef4f44">如果采用前面推导中<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的总响应约定，那么就是<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。关键点不在这个因子，而在于：磁化率正比于总磁矩涨落。</div><div class="notion-text notion-block-39d3936163f481df90b1f873ccd902aa">再定义连通关联函数</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4815095bafd434d983768">因为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4810aadb9c3d8d0f492c9">所以在平移不变体系中</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4818cacdfe6e5769ecd0d">这说明磁化率不是一个孤立的响应系数，而是在空间上把所有自旋之间的关联加起来。远离临界点时，关联函数通常近似指数衰减：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4816b8ea4f73609e9a901">这里<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就是关联长度，表示一个自旋取向的影响大约能传到多远。</div><div class="notion-text notion-block-39d3936163f481e1b606ead705d78745">从团簇图像看，系统可以想成由许多局域上较一致的自旋团簇组成。高温时团簇很小，向上团簇和向下团簇频繁打散，外场只能影响很短距离内的自旋，因此<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 有限，磁化率也有限。低温深处已经形成宏观有序背景，涨落同样不强。</div><div class="notion-text notion-block-39d3936163f481528262fbbc2ea5924c">接近临界温度时，不同尺度的团簇同时出现，小团簇嵌套在大团簇中，系统不再有一个固定的典型长度。此时关联长度迅速增大，并在热力学极限下满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812db640c1625ef212df">当<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，关联函数的空间积分变得越来越大：一个很小的外场不再只推动少数自旋，而是能让一个很大的相关区域一起偏向同一方向。于是磁化涨落增强，磁化率出现临界发散：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481f6812bd1baf15ddc0f">所以所谓“物理量发散”，并不是说某个数学符号失去意义，而是说在热力学极限中，系统对外界扰动的响应没有普通有限尺度可以截断；临界点附近的集体涨落变成了宏观现象。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f481979a1ef1627f6616a0" data-id="39d3936163f481979a1ef1627f6616a0"><span><div id="39d3936163f481979a1ef1627f6616a0" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f481979a1ef1627f6616a0" title="7. 平均场理论的物理意义与局限"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">7. 平均场理论的物理意义与局限</span></span></h3><div class="notion-text notion-block-39d3936163f481fbb94fe326f1482703">平均场理论的优点是抓住了相变的骨架：配分函数、自由能、序参量、自洽方程、临界温度、对称性破缺和幂律奇异性。可是这里立刻出现一个问题：如果临界点附近关联长度<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 发散，平均场近似会被长程涨落严重干扰，那么为什么临界点附近的幂律行为仍然有效？幂律行为难道不是平均场失效造成的吗？</div><div class="notion-text notion-block-39d3936163f481629fc2f29fedacf6d5">答案是：平均场失效，主要是因为它低估了空间涨落，从而给不准临界指数；但幂律本身不是平均场的产物，而是临界点无特征尺度的结果。</div><div class="notion-text notion-block-39d3936163f481498b05d5f6390ce7b7">先看正常情形。远离临界点时，连通关联函数通常指数衰减：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481ccbb05f2f2269b324f">这里<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是特征尺度：当距离<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 达到几个<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 以后，一个自旋对另一个自旋的影响就迅速消失。系统能够“记住”一个典型长度，因此响应函数一般是有限的。</div><div class="notion-text notion-block-39d3936163f48118a88cd50c6c34cfdc">临界点不同。此时<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，指数衰减中的固定长度尺度消失了。没有特征尺度，系统就不能“记住”某个固定长度、固定能量或固定时间；把观察尺度放大以后，系统的统计形态仍然相似。这叫尺度不变性，或者自相似性。</div><div class="notion-text notion-block-39d3936163f4816bbc6ee84b55a2c4c6">用数学语言说，若某个物理量<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在尺度变换<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 下只能改变一个整体倍数，就应满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48131b061e854dff2500f">这种缩放方程的解就是幂律：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481aaa0d0cebd1a95e27f">所以在临界点，关联函数不再以某个固定长度截断，而通常变成幂律衰减：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48189b671fbbc6b04f99e">同样，关联长度、磁化率和自发磁化也在临界点附近表现为幂律：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e281dde5df718e75f1">因此要区分两件事：幂律形式来自临界点的无尺度性；临界指数的具体数值来自涨落如何在不同尺度之间耦合。平均场近似的问题不是“推出了幂律”，而是把空间涨落平均掉了，所以通常只能给出错误的指数。对于 Ising 模型，平均场给出<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，这些结果在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时可靠；在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，涨落会改写指数，但不会消灭临界幂律本身。</div><h2 class="notion-h notion-h1 notion-default notion-h-indent-0 notion-block-39d3936163f4815ab280d74bddb578ad" data-id="39d3936163f4815ab280d74bddb578ad"><span><div id="39d3936163f4815ab280d74bddb578ad" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4815ab280d74bddb578ad" title="Some Exact Results for the Ising Model"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Some Exact Results for the Ising Model</span></span></h2><h3 class="notion-h notion-h2 notion-default notion-h-indent-1 notion-block-39d3936163f481afbf95f5f7cace11ba" data-id="39d3936163f481afbf95f5f7cace11ba"><span><div id="39d3936163f481afbf95f5f7cace11ba" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f481afbf95f5f7cace11ba" title="1. 一维 Ising Model 的精确解"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. 一维 Ising Model 的精确解</span></span></h3><div class="notion-text notion-block-39d3936163f48191853beb247048c604">一维 Ising 模型是最简单的可精确求解模型。设有<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个格点，并取周期边界条件<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，哈密顿量为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481849bc8ee11039ace9b">配分函数为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4819a8848d16442df5aa7">为了把它写成相邻自旋之间的矩阵乘积，把外场项平均分配到相邻两条键上：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481b28a88dfe22e19adbe">定义转移矩阵</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481099a1ec009318960e2">配分函数是对所有自旋构型求和：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481978665c0126b6319e3">显式写成矩阵就是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a082f9c78793bb1f21">周期边界条件使配分函数变成矩阵迹：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481ae81f2d4a3a02cbf9e">转移矩阵的两个本征值为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481cd8eccdd5830492914">因此精确配分函数是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481f99b2efff940fc22be">在热力学极限<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中，由最大本征值控制：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4810cb4b1c66a380d9eca">磁化强度由自由能对外场的导数给出：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a0ba72da1f3b62d65d">直接计算得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481d682eef602292f6f47">于是当<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 且温度有限时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a4a552febd6bc1e157">这说明一维 Ising 模型在任何有限温度下都没有自发磁化，也没有有限温度相变。只有在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，系统才完全有序。</div><div class="notion-text notion-block-39d3936163f4811e930cc2a9b4da6eaa">这个结论也可以从关联长度看出。在零外场<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，本征值化为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4814d997fe2c9a94a093c">一维关联函数为</div><div class="notion-text notion-block-39d3936163f4818f9338e85261cd0a27">从关联函数的热平均定义出发，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48129b189fb514fa65877">定义自旋算符矩阵</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481b8b9a9c77fdc6b1fb7">在转移矩阵乘积中插入一个<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就等价于给该格点乘上对应的自旋值。因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4816f8fd6f47f630b4d06">在零外场时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的两个归一化本征态可以取为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48168996fce74e28e45f6">对应本征值<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。自旋算符会把这两个本征态互相变换：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48142a8b8c50dd6f708c9">因此在本征态基底中，迹包含两条闭合传播通道。第一条从<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 出发：经过<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 得到因子<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，插入<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 后变成<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，再经过<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 得到因子<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，最后第二个<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 把它带回<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这给出一项<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-39d3936163f48182bf1fdbcbe4287620">第二条从<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 出发，同理得到另一项<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。所以分子更准确地是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481389b44c802f3f98002">分母为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481748abed675ae737f8b">由于零场有限温度下<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 且<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 固定时，第二条通道<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 相对第一条通道被因子<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 压低，因此可忽略。于是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4817ba3c1f51722a0dd96">写成指数衰减形式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481aa9ab9e70444d952fa">得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48147bdd4fd760688038f">只要<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，就有<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 有限；只有<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，关联长度才发散。</div><div class="notion-text notion-block-39d3936163f481beb2d3c1ef0a09b365">这个结果也可以用畴壁来理解。所谓畴，是指一段局域上自旋取向一致的区域；畴壁就是两个取向不同的畴之间的边界。例如一维链中</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-default notion-block-39d3936163f48150b792e30029c1905e">中间的反向小畴两端各有一个畴壁：左端是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，右端是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。畴壁会把原本的长程同向排列切成不同取向的片段。</div><div class="notion-text notion-block-39d3936163f4815d9d59ee1c068259ed">畴壁的能量代价可以直接从单条键的能量看出。相邻自旋如果同向：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48136964dd052db645f64">这一条键的能量是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48111bde6f585bae5a0c6">如果相邻自旋反向：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481028583e3bde003ab3c">这一条键的能量是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481c98920f363fd81a38a">所以一个畴壁相对于同向排列多出来的能量是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48172bd11dec27ed9552e">物理上，一维系统中一个畴壁只需要有限能量<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在任何有限温度下，畴壁都有非零热概率；热力学极限中会出现有限密度的畴壁，把长程序打断。因此一维短程 Ising 模型没有有限温度铁磁相变。</div><div class="notion-blank notion-block-39d3936163f4801ba211cd712d155000"> </div><h3 class="notion-h notion-h2 notion-default notion-h-indent-1 notion-block-39d3936163f481efb7f5f954ca148de8" data-id="39d3936163f481efb7f5f954ca148de8"><span><div id="39d3936163f481efb7f5f954ca148de8" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f481efb7f5f954ca148de8" title="2. 二维 Ising Model：低温与 Peierls 液滴"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. 二维 Ising Model：低温与 Peierls 液滴</span></span></h3><div class="notion-text notion-block-39d3936163f481459231f2fa3cf819dd">二维最近邻 Ising 模型定义在平方格子上：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481eab9f0f62f27496266">低温展开的出发点是：当<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，配分函数主要由两个完全有序基态贡献：全向上和全向下。低能激发则可以按“翻转少量自旋造成的畴壁能量代价”来系统计数。</div><div class="notion-text notion-block-39d3936163f4814ba187d86b62aa7588">设平方格子有<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个格点，周期边界下最近邻键数为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在全有序基态中，每条键都是同向的，基态能量为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a7a159ef18368b7b93">因此两个基态给出因子</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481ab8adcd6ac3a182287">接下来考虑翻转自旋产生的低能激发。一个孤立反转自旋有四条边界键，每条边界键把同向键变成反向键，能量从<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 变成<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以每条边界键多付出<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。于是孤立反转的能量代价是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4813380eece9fed632729">有<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置，所以给出<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。两个相邻反转自旋形成一个长度为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的边界，能量代价为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，相邻格点对数为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以给出<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。于是低温展开开头为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4816e848ed8d5194beec5">这串展开的意义是：低温涨落可以组织为反向自旋区域的边界。一个由负自旋组成的连通区域称为低温背景中的液滴；液滴外侧的闭合边界称为 Peierls 轮廓。若轮廓长度为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48181a046ff6dec61ee6f">因此单个长度为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的轮廓带来 Boltzmann 权重</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48163bf7ac190fd6e0967">但是轮廓的形状数也会随<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 增长。粗略地，把长度为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的可能轮廓数写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481bcb130e4bad60e2aae">这里<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是轮廓熵密度，表示单位边界长度大约带来多少形状选择。于是所有长度的液滴贡献可以估计为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48165a518e29d66ed50fb">这就是二维低温相变的能量-熵竞争。能量项<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 压制长边界液滴；熵项<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 鼓励更多形状。低温时<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，大液滴被压制，有序相稳定；升温后<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，熵增压过边界能，大液滴大量出现，有序相被破坏。粗略临界条件为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e585d9f99612aca2c2">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4819f8fc0c16f88335a52">这个式子不是精确临界温度，而是解释为什么二维会有有限温度相变的量级估计。二维 Ising 的精确结果是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481798dc1e1c49f8435fc">一维为什么不发生有限温度相变，也可以从同一个能量-熵竞争看出。一维畴壁是点，一个畴壁的能量代价只是<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，但它可以放在<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置上，因此自由能代价近似为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481c08abec69cc24d37e6">在热力学极限<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中，只要<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，熵项最终都会压过有限能量代价，畴壁不可避免，长程序被切断。二维中则不同：一个液滴的代价随边界长度<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 增长，低温下可以压过形状熵，因此有序相能够稳定。</div><div class="notion-text notion-block-39d3936163f4812e89e8cb31ebb65330">最后区分一下液滴和团簇。液滴通常是在低温有序背景中出现的少数相区域，例如全向上背景中的一块负自旋区域；它强调的是边界能量和低温激发。团簇则更一般，指相同取向或强相关自旋组成的连通区域，尤其在临界点附近会出现多尺度分布。低温液滴可以看成一种特殊团簇，但临界团簇不一定嵌在稳定背景中，也不一定只是少数相液滴；它们更强调关联长度发散和无尺度涨落。</div><h3 class="notion-h notion-h2 notion-default notion-h-indent-1 notion-block-39d3936163f481d2ad59f7824c05b43f" data-id="39d3936163f481d2ad59f7824c05b43f"><span><div id="39d3936163f481d2ad59f7824c05b43f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f481d2ad59f7824c05b43f" title="3. 二维 Ising Model：高温"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 二维 Ising Model：高温</span></span></h3><div class="notion-text notion-block-39d3936163f48106bf3cf13b3170d1d7">低温展开从全有序基态出发，把涨落组织成反向液滴的边界。高温展开则从无序相出发，此时<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 较小，邻近自旋之间的相关性弱。为简单起见，先取零外场<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。配分函数为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a1b191da305aa789fe">对每一条最近邻键使用恒等式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481ff91b8f64a0c6c8167">这是因为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。若记<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并设最近邻键总数为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481458326de553eef2e32">现在展开乘积。对每条边，或者选择<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，或者选择<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此展开中的每一项都对应格子上一组选中的边<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，其权重为<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并带有自旋乘积。</div><div class="notion-text notion-block-39d3936163f481db8aa5dbbcc37e1c7e">对自旋求和时，只有每个格点连接偶数条被选中边的图形会保留下来。原因是如果某个格点<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在该项中出现了奇数次，则有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481afb899ee0aa92f6330">例如，展开时当然可能选中两条彼此不相连的边<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，这一项为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a0bdd3f454aaa2d7f7">但是在配分函数中还要对所有自旋求和。由于</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4815a8d94d1cadbed57ff">只要某个自旋在该项中出现奇数次，对这个自旋的求和就会把整项消掉。上面的例子里<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 都只出现一次，因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481aca32fc29456133453">再看两条相邻边<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，对应项为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4810987d4c3672e7ad6f4">虽然中间的<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 出现两次并变成<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，但端点<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 仍然各出现一次，所以这项也被求和消掉。</div><div class="notion-text notion-block-39d3936163f481949906cfb6a4333ef0">相反，如果选中一个闭合小方格的四条边，就得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4818280d4e4710e8b3a73">每个顶点都出现两次，因此对自旋求和不会为零。这就是为什么高温展开最后只留下闭合回路，或者更一般地说，只留下每个顶点度数为偶数的图形。</div><div class="notion-text notion-block-39d3936163f48104ab31d1c33b6c0cfe">只有<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 为偶数时，求和才给出<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。所以高温展开留下来的图形是闭合回路，或者若干个互不相交/相交于偶数度顶点的闭合图形。最终得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4810c9252f5a76fd3d85f">这个公式就是二维 Ising 模型的高温展开。高温时<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。长回路权重<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 很小，说明大尺度相关被强烈压制，系统处在无序相。</div><div class="notion-text notion-block-39d3936163f4816abf20c3d78dd446a5">这和低温展开形成互补。低温展开中的基本对象是反向液滴的边界，权重形如<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；高温展开中的基本对象是闭合回路，权重形如<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这两个结构在对偶格子上具有相同的图形形式。</div><h3 class="notion-h notion-h2 notion-default notion-h-indent-1 notion-block-39d3936163f4810d8a12c405a8080544" data-id="39d3936163f4810d8a12c405a8080544"><span><div id="39d3936163f4810d8a12c405a8080544" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4810d8a12c405a8080544" title="4. Kramers-Wannier 对偶性"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. Kramers-Wannier 对偶性</span></span></h3><div class="notion-text notion-block-39d3936163f481879c41e43d2edd99f7">因此高温展开自然引向 Kramers-Wannier 对偶性。若令对偶耦合<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48116b3abf1262ce0b672">就可以把高温展开映射为对偶格子上的低温展开。等价地，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481b69325d6fc2908582f">这里要特别注意：对偶性不是把原格点上的自旋<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 一一变成同一批实际点上的自旋。高温展开先在原格子上对自旋求和，求和之后留下的对象不是自旋构型，而是闭合边集合<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。低温展开也把自旋构型重新编码成畴壁轮廓。因此对偶性比较的是展开后的闭合线图形，而不是原始自旋点本身。</div><div class="notion-text notion-block-39d3936163f48150a9cde0b4d6bae087">长度<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 也不是欧氏长度，而是边的条数。每一条原格子边都和一条对偶格子边相交，因此原格子上选中<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 条边，等价于对偶描述中对应<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 条边。正是这个一一对应使得<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可以和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 逐项比较。</div><div class="notion-text notion-block-39d3936163f4819f82e3e2d355f29805">更严格地说，对偶性来自比较同一组闭合线在高温展开和对偶低温展开中的权重。原格子的高温展开为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-default notion-block-39d3936163f48165aef5ee0fd9ed71c5">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是原格子高温展开中留下的闭合边集合。每一条原格子边都唯一对应一条穿过它的对偶边；反过来，在对偶模型的低温展开中，畴壁也可以用穿过对偶键的原格子边来表示。因此两边都可以用同一类闭合边集合 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 来标记。这里比较的是边集合的长度和权重，而不是把原格点自旋直接变成对偶格点自旋。</div><div class="notion-text notion-default notion-block-39d3936163f481b0a42cfbcc6a548b09">对偶模型的自旋定义在对偶格点上，耦合为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。它的低温展开从两个有序基态出发；对偶自旋的正负畴之间形成闭合畴壁。把这些畴壁重新画回原格子的边集合后，若对偶格子有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个格点、<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 条边，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481958283cf1b468db6eb">两边的闭合图形求和完全相同。为了让每条轮廓的权重逐项对应，必须要求</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4817ebdb6c5b8d6ba830b">于是闭合图形的求和部分相同，只差前面的解析因子。也就是说</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4815399aded2390b655d0">这里并不是把前面的系数丢掉了。逐项比较闭合线时，我们只需要匹配每单位长度的权重<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；但完整的配分函数关系必须保留<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这些前因子给自由能贡献解析的背景项，而临界奇异性来自闭合线求和中大回路或大畴壁的统计。</div><div class="notion-text notion-block-39d3936163f481f48255ea8fefea9b02">这就是配分函数层面的 Kramers-Wannier 对偶关系。它说明原模型在耦合<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的高温展开，等价于对偶模型在耦合<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的低温展开，除了前面那些不会导致奇异性的解析因子。</div><div class="notion-text notion-block-39d3936163f481b989d2e08e194a45c6">再把<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 改写成对称形式。令<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48190b463eddbbf9ba0a5">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481499235e0c0d4cf4f0e">相乘得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481c6af43d94ed4a6a3eb">对于平方格子，原格子和对偶格子仍然是平方格子，所以若相变点唯一，临界点必须落在自对偶点<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481fba45ec6606f76adeb">从而得到精确临界温度</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h3 class="notion-h notion-h2 notion-default notion-h-indent-1 notion-block-39d3936163f48120a498e5d9aadbc93f" data-id="39d3936163f48120a498e5d9aadbc93f"><span><div id="39d3936163f48120a498e5d9aadbc93f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f48120a498e5d9aadbc93f" title="5. van der Waals 方程与 Landau 理论的动机"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. van der Waals 方程与 Landau 理论的动机</span></span></h3><div class="notion-text notion-block-39d3936163f481a9992bf47e0e77c6a8">现在可以回头理解一个重要事实：van der Waals 方程和平均场 Ising 模型虽然描述的微观对象完全不同，一个是流体的液-气转变，一个是晶格自旋的铁磁转变，但它们在临界点附近给出同一组平均场临界指数。</div><div class="notion-text notion-block-39d3936163f481109becd07815bbf92b">这里不要把 van der Waals 方程和 Clapeyron 方程混在一起。Clapeyron 方程</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481aa8014c45db3406642">描述的是一级相变共存线的斜率；而 van der Waals 方程是一个具体的平均场状态方程，能够描述液-气共存线终止处的临界点：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812483bedd304d5e9a17">在液-气临界点附近，可以把序参量取为密度相对于临界密度的偏离，例如</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481cb93e1c7b0c7bf9ca1">在 Ising 模型中，序参量是磁化强度<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。两者的物理含义不同，但在临界点附近都可以用同一个形式的有效自由能描述：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481d998cbfeaaca9a0c9c">这就是 Landau 理论的核心工作：不再从具体微观模型出发，而是写下满足对称性和稳定性要求的序参量自由能。对于 Ising 铁磁体，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，外场<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是磁场；对于液-气系统，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，与其共轭的场可以理解为化学势偏离临界值。</div><div class="notion-text notion-block-39d3936163f48138b2cbc047b1c9a3af">从这个自由能出发，极小化条件给出平均场临界行为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4814b87d8c212b1ee77c5">也就是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812bbe0bd4c90dce78a2">这正是 David Tong 所说的：van der Waals 方程和平均场 Ising 模型给出相同的临界指数。它们相同，不是因为液体分子真的等同于 Ising 自旋，而是因为二者在临界点附近被同一个平均场序参量自由能控制。</div><div class="notion-text notion-block-39d3936163f481e5a908e2beff9a912d">这些答案有时是错的，是因为平均场自由能忽略了临界涨落。真实三维液-气临界点属于三维 Ising 普适类，临界指数并不等于平均场值。Landau 理论的价值在于统一了不同系统的平均场结构；它的局限则提示我们下一步必须加入空间涨落，也就是 Landau-Ginzburg 理论。</div><div class="notion-blank notion-block-39d3936163f480ee9e50d7f1d36162fd"> </div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-39d3936163f48142be8bf40181bdaef1" data-id="39d3936163f48142be8bf40181bdaef1"><span><div id="39d3936163f48142be8bf40181bdaef1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f48142be8bf40181bdaef1" title="Landau Theory and Landau-Ginzburg Theory"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Landau Theory and Landau-Ginzburg Theory</span></span></h2><div class="notion-text notion-block-39d3936163f4812681feea1f7bd88273">前面我们从 Ising 模型和 van der Waals 方程看到一个共同结构：相变附近最重要的变量不是所有微观自由度，而是一个宏观序参量。Landau 理论的出发点就是把自由能直接写成序参量的函数，然后用对称性和稳定性限制它的形式。</div><div class="notion-text notion-block-39d3936163f481e29764fae9705e3741">因此 Landau 理论不是从配分函数精确求和开始，而是把很多微观细节压缩进少数系数中。它回答的问题是：如果我知道相变的序参量、对称性和稳定性，那么自由能景观会怎样改变？</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f4813e984bea24f4a21977" data-id="39d3936163f4813e984bea24f4a21977"><span><div id="39d3936163f4813e984bea24f4a21977" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4813e984bea24f4a21977" title="1. Landau 理论的出发点"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Landau 理论的出发点</span></span></h3><div class="notion-text notion-block-39d3936163f481c6a314e47fe29d7030">设序参量为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在铁磁体中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；在液-气相变中，可以取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。Landau 理论假设，在临界点附近自由能密度可以按序参量展开：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4819585fac74cb9bcd61f">如果体系有 Ising 型反演对称性 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，且外场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，自由能必须在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 下不变，因此奇次项消失。外场是和序参量共轭的变量，会加入一项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。最常用的 Landau 自由能写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48198ad52c2a3344378cb">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是控制参数，随温度穿过零；<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 保证大 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时自由能向上，不会无界下降。平衡态不是由 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 单独决定，而是由自由能极小值决定：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f4813b9c78f0cbed2fcebc" data-id="39d3936163f4813b9c78f0cbed2fcebc"><span><div id="39d3936163f4813b9c78f0cbed2fcebc" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4813b9c78f0cbed2fcebc" title="2. 二级相变"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. 二级相变</span></span></h3><div class="notion-text notion-block-39d3936163f48128b399c35c62fd933f">先考虑 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 且 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的情形：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481a788a8f3e7c4d006c0">极值条件为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812ba8bef2c107e25713">所以可能的解是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4819b8552f087416adff4">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，只有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是稳定极小值，体系处在无序相。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481f8876ffdbd6df153c4">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 变成极大值，而两个新的极小值出现：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812ab752ffc72ea38346">序参量从零连续长出来，因此这是一种连续相变，也叫二级相变。自由能景观不是突然换成另一张图，而是随着 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 穿过零，中心极小值逐渐变平、失稳，再分裂成两个对称极小值。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39d3936163f48139b434e99c3bedacc5"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://www.notion.so/image/attachment%3A89acab74-a697-4c28-bf08-97e41837445c%3Alandau_second_order_free_energy.png?table=block&amp;id=39d39361-63f4-8139-b434-e99c3bedacc5&amp;t=39d39361-63f4-8139-b434-e99c3bedacc5" alt="Landau free-energy landscape for a continuous transition: above Tc the stable minimum is at φ = 0; below Tc the center becomes unstable and two symmetry-related minima appear." loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">Landau free-energy landscape for a continuous transition: above Tc the stable minimum is at φ = 0; below Tc the center becomes unstable and two symmetry-related minima appear.</figcaption></div></figure><div class="notion-text notion-block-39d3936163f481a98ad9e3a210212f3c">外场不为零时，状态方程为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48174a081d12ac5370938">在高温侧 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 且小外场下，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48168a961f75fec225ea3">在临界等温线 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，于是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481268642e1c23b20e2eb">因此 Landau 理论给出平均场指数</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f481ceaaa9c118d87ab28e" data-id="39d3936163f481ceaaa9c118d87ab28e"><span><div id="39d3936163f481ceaaa9c118d87ab28e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f481ceaaa9c118d87ab28e" title="3. 一级相变"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 一级相变</span></span></h3><div class="notion-text notion-block-39d3936163f4814da458c782ecf6f76e">二级相变的典型 Landau 图像依赖 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对称性，因此自由能只含偶次项。一级相变可以从另一种更直观的情形看出来：如果自由能中允许奇次项，或者 Ising 模型中加入外场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，那么左右两个方向不再等价。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48153812ac0dd24083bc4">只要存在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 这种奇次项，就有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4814b826cff1304357257">这不是自发破缺，而是显式破坏对称性：外场已经偏好某一个符号的磁化。平均场 Ising 自由能在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时就是这种情况。把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 展开后会出现 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的幂，因此展开成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的函数时会混入奇次项。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481cb84f9e8180eab6a89">低温时自由能仍然可能有双井结构，但因为对称性已经被外场破坏，两个极小值通常不一样深。较低的极小值是真正的平衡态；另一个较高但仍然局域稳定的极小值是亚稳态。</div><div class="notion-text notion-block-39d3936163f4810bbff9e8a45c364549">亚稳态为什么能暂时存在？因为它虽然不是全局最低，但仍然满足局域稳定条件：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e4a78ac427a58f39b3">系统如果已经落在亚稳态井里，不能通过无穷小变化直接滑到真正稳定态，而必须先越过中间的自由能势垒。因此一级相变常伴随过冷、过热、成核和滞后现象。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39d3936163f481c4a15cc58b4fbede42"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://www.notion.so/image/attachment%3Ad916831c-3cb7-4b39-b63e-3d885ec905ca%3Alandau_first_order_free_energy.png?table=block&amp;id=39d39361-63f4-81c4-a15c-c58b4fbede42&amp;t=39d39361-63f4-81c4-a15c-c58b4fbede42" alt="First-order transition as exchange of the global minimum: one side stable, coexistence with equal depths, and the other side stable." loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">First-order transition as exchange of the global minimum: one side stable, coexistence with equal depths, and the other side stable.</figcaption></div></figure><div class="notion-text notion-block-39d3936163f481fd8cdbf2a08d23c3c0">一级相变点不是某个极小值刚刚失稳的位置，而是两个局域极小值的自由能相等的位置。若两个相对应的序参量为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，相变点满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e18c5eeff312c2d393">跨过这个点时，全局最低点从一个井切换到另一个井，因此平衡序参量发生有限跳变：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481be9f8dfa203e5c4dea">这就是一级相变的核心：不是序参量从零连续长出来，而是两个局域稳定相竞争，最终由自由能较低者成为真正平衡态；在共存点两个相自由能相同，跨过共存点序参量跳变。液-气相变的一阶共存线就是这种图像：气相极小值和液相极小值在共存线上自由能相等。</div><div class="notion-text notion-block-39d3936163f481d6b01fcf4fa7a9282e">亚稳态最终消失的位置叫 spinodal。它不是相变点，而是局域极小值和中间极大值合并、势垒消失的位置，由</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4812f91fcfd4cc1ecad7e">决定。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39d3936163f4815e8291db3417306cda" data-id="39d3936163f4815e8291db3417306cda"><span><div id="39d3936163f4815e8291db3417306cda" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39d3936163f4815e8291db3417306cda" title="4. Landau-Ginzburg Theory"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. Landau-Ginzburg Theory</span></span></h3><div class="notion-text notion-block-39d3936163f48147aef5ceef662f9447">Landau 理论把序参量当成空间均匀的数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。但临界点附近最重要的正是长波长空间涨落，因此必须把序参量升级为空间场：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f480b9b992e4926a36bb4f">但是Ising模型中最小尺度是格点，因此<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>实际上是粗例化后局域平均。从而Landau-Ginzburg 自由能泛函写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481548f92d3138aef6feb">这里局域势能项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 决定每一点更喜欢无序还是有序；梯度项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 惩罚空间变化，表示相邻区域的序参量不愿意突然跳变。</div><div class="notion-text notion-block-39d3936163f481a2959bf7be012c0661">若只看高温无序相中的小涨落，可以暂时忽略四次项，得到高斯近似：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48112a924c83aa938df1b">下面把这个式子变到傅里叶空间。约定</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48174bbb0e3b06685359a">如果 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是实场，则傅里叶分量满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f48172ac78ceabec2a48fe">先看质量项：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4814c8776d61d1772aa81">利用空间积分给出的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 函数</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481f1a81cc4ed8a0bc454">于是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，从而</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4811381aced9b32fac851">再看梯度项。傅里叶空间中求导等价于乘以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4818fad32edaeea59ca3d">同样利用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481cc9005f97a0c8ba63d">也就是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f481e0aa17c87f366477b7">于是动量空间关联函数具有 Ornstein-Zernike 形式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39d3936163f4816a80e6c10b7ecd67d0">把分母写成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，可以读出关联长度</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ab94bef4cd00072f01">这个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 之所以叫关联长度，是因为动量空间中的这种形式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b8ab20d7af57208d17">傅里叶变回实空间后会给出指数衰减。也就是说，实空间关联函数大致满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b2a761cd212215bf57">更精确地说，在高维中还会乘上一个随距离缓慢变化的幂次因子；但控制关联在多远距离上被截断的主要尺度就是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。当 Landau 系数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，说明长距离关联不再被有限尺度截断。</div><div class="notion-text notion-block-39d3936163f481f39da1c5c70f349ad0">这说明 Landau-Ginzburg 理论已经开始描述空间关联：当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，长波长模式 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的代价变小，关联长度发散。</div><div class="notion-text notion-block-39d3936163f4810a87ddc6738e2b6583">不过高斯近似仍然是平均场性质的。真正的问题是四次相互作用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 会让不同波长的涨落彼此耦合。这个耦合在高维不重要，平均场指数可用；在低维和临界点附近会变得重要，从而修正临界指数。这就是从 Landau-Ginzburg 理论走向重整化群的入口。</div><div class="notion-text notion-block-39e3936163f4812ea925d39111e6290b">因此，高斯近似虽然已经允许 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在空间中涨落，但它仍然给出和 Landau 理论相同的平均场结果。原因是：高斯近似只保留二次项，每个傅里叶模式都是独立的高斯变量，涨落模式之间没有相互作用。</div><div class="notion-text notion-block-39e3936163f48186b645ff53baa69000">例如在高温侧，均匀 Landau 理论在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近展开到二次阶就是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4810a8003efccc2d67bdb">Landau-Ginzburg 高斯近似只是把这个二次涨落推广到所有波长：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f982eed6cb0ab489c2">所以它能给出空间关联和关联长度，但临界指数仍然是平均场的。例如</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481cda168f6c82a50a5e7">真正超出平均场的是四阶项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在傅里叶空间中，它不再是单个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 模式的平方，而是把四个模式耦合起来：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a4be78e6d606710a39">临界点附近 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，长波模式变软，涨落变大，模式之间的耦合就可能不能忽略。平均场指数是否被修正，取决于这些涨落相互作用在长距离下是否重要；这正是 Wilson 重整化群要回答的问题。</div><div class="notion-blank notion-block-39d3936163f480a7ab80e27fc85748a0"> </div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Numerical Renormalization Group]]></title>
            <link>https://blog.xiangsiqi.site/39f39361-63f4-809e-a6cd-fff49ec19581</link>
            <guid>https://blog.xiangsiqi.site/39f39361-63f4-809e-a6cd-fff49ec19581</guid>
            <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39f3936163f4809ea6cdfff49ec19581"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39f3936163f48149b66bef968e39435a" data-id="39f3936163f48149b66bef968e39435a"><span><div id="39f3936163f48149b66bef968e39435a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f48149b66bef968e39435a" title="1. Ising 模型的数值采样：典型方法与主要问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Ising 模型的数值采样：典型方法与主要问题</span></span></h3><div class="notion-text notion-block-39f3936163f481459344e72ec410a432">在 Kadanoff 数值重整化群中，第一步通常是从原始 Ising 模型产生可靠的微观系综。后续的 blocking、有效耦合反演以及临界指数估计，都依赖于该系综是否正确代表目标 Boltzmann 分布。因此，采样算法的选择、热化判据、自相关时间估计和误差分析构成数值 RG 的基础环节。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481d083b3eb595385065d" data-id="39f3936163f481d083b3eb595385065d"><span><div id="39f3936163f481d083b3eb595385065d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481d083b3eb595385065d" title="1.1 Ising 模型与目标分布"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1.1 Ising 模型与目标分布</span></span></h4><div class="notion-text notion-block-39f3936163f481c5b440ef2a973a3c54">以二维方格 Ising 模型为例，微观变量为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，哈密顿量可写为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481f393fbeaabbe0b9050">数值采样的目标是在给定 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 下产生满足 Boltzmann 分布的构型序列：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481509bbfef8c3f3262e7">后续所有热平均均由样本平均近似，例如</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481198ae6fba75ed6a2ac" data-id="39f3936163f481198ae6fba75ed6a2ac"><span><div id="39f3936163f481198ae6fba75ed6a2ac" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481198ae6fba75ed6a2ac" title="1.2 局域 Metropolis 更新（1953）"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1.2 局域 Metropolis 更新（1953）</span></span></h4><div class="notion-text notion-block-39f3936163f4816298a8ef793af1d298">最基本的 Markov Chain Monte Carlo 方法是单自旋翻转 Metropolis 算法。每一步随机选择格点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，提出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并按接受率</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481fc981eedccf7426f08">接受或拒绝。该方法实现简单，且容易检验 detailed balance；但在临界点附近，由于相关长度增大，单自旋局域更新难以快速改变大尺度畴结构，因此会出现严重的临界慢化。</div><div class="notion-text notion-block-3a03936163f4814eb41bdcc0602fa303">该接受率的有效性来自 detailed balance。设目标平衡分布为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f48122962ce9726f5d5fd5">若 proposal 概率关于正反过程对称，即 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则 detailed balance 条件</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f481a19612e2bb9bee3c85">等价于要求接受率满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f481d6925fd45da3918f73">Metropolis 选择</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f481a3b239e081c1afb075">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，新构型能量较低，正向过程必然接受；反向过程的接受率为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，正向过程以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 接受，反向过程必然接受。两种情况下均有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f481f39845c0f42467d5e5">因此 Boltzmann 分布是该 Markov 链的不变分布。只要链同时满足遍历性，长时间采样即收敛到目标热平衡分布。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481468105ea4e2f1e0c07" data-id="39f3936163f481468105ea4e2f1e0c07"><span><div id="39f3936163f481468105ea4e2f1e0c07" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481468105ea4e2f1e0c07" title="1.3 Cluster 更新：Swendsen-Wang（1987）与 Wolff（1989）"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1.3 Cluster 更新：Swendsen-Wang（1987）与 Wolff（1989）</span></span></h4><div class="notion-text notion-block-39f3936163f48133bf9ef9870e80f082">对铁磁 Ising 模型，临界点附近更常用 Wolff 或 Swendsen-Wang cluster 算法。其基本思想是按照 Fortuin-Kasteleyn 表示构造同向自旋簇，并整体翻转一个或多个簇。相邻同向自旋之间的成键概率为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f481c08a0bc5424e48d21f">FK 表示的出发点是将每条边的 Boltzmann 权重改写为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f4816ca7bace52e0fabe2a">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，上式给出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，上式给出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此该改写与原 Ising 权重完全等价。</div><div class="notion-text notion-block-3a03936163f48118b8bdd516c7e4542f">若目的仅是证明配分函数等价，上述单边恒等式已经足够。引入 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的作用不是改变模型，而是把括号中的两项解释为一个可采样的二值辅助变量。具体地，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f48106aea9ddf441ddb724">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对应“不放置 bond”，贡献 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对应“放置 bond”，贡献 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此只有同向自旋之间才能被连接。</div><div class="notion-text notion-block-3a03936163f4814e94e8e6886a340ad6">这种扩展变量的价值在于把代数分解转化为联合概率模型 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在该扩展空间中，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 只是“同向边以概率 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 连边”，而 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 只是“每个 cluster 独立赋值为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>”。因此，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是使 cluster update 成为严格 Gibbs sampling 的辅助变量。对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 求和后仍恢复原来的 Ising 自旋分布。</div><div class="notion-text notion-block-3a03936163f481e5b169d05d19a7543e">引入 bond 变量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 后，可定义联合分布</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f4811890b5fed55cf9c4f1">该式说明：异号相邻自旋之间不能放置 bond；同号相邻自旋之间以概率 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 放置 bond。对 bond 变量求和后恢复原 Ising 权重，因此该联合表示在边缘分布意义下等价于原自旋模型。</div><div class="notion-text notion-block-3a03936163f4816cbbfcffde68bd8a0b">更明确地说，该联合分布的归一化常数与原 Ising 配分函数只差一个由单边恒等式带来的整体因子。对所有边相乘并在扩展空间中求和，有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f48157a0b4fd653b01e2ee">因此，若记上式乘积为扩展空间中的未归一化权重 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则联合概率分布为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这一步只说明 FK 表示与原模型等价；要得到采样算法，还需要考察该联合分布的条件分布。</div><div class="notion-text notion-block-3a03936163f4810d90ecdd22f354a146">首先固定自旋构型 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是逐边乘积，且在给定 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 后每条边的 bond 变量只出现在对应边因子中，故条件分布因子化为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f4812cb963e546034d17a8">对单条边，有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f48142831ccd1301ea059b">若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，只能取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则该边权重为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，归一化后得到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这正是 cluster 算法中“只在同向邻居之间以概率 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 放置 bond”的规则。</div><div class="notion-text notion-block-3a03936163f481d3b355c328f089510c">其次固定 bond 构型 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，采样自旋构型 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。联合权重中的因子 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表明，只要某条边被占据，即 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，就必须有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此，同一个 connected cluster 内的所有自旋必须相同，而不同 clusters 之间没有相互约束。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a03936163f481e1b651f0f1751fc737">对 Ising 模型，每个 cluster 因而可以独立赋值为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，且两者权重相同，故 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。于是，交替采样 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就给出 Swendsen-Wang 算法：先由当前自旋构型生成 FK bonds，再对每个 cluster 独立重采样自旋。</div><div class="notion-text notion-block-3a03936163f481419b71da4cca783681">Swendsen-Wang 算法可视为对联合分布 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 Gibbs sampling，其步骤为：</div><ol start="1" class="notion-list notion-list-numbered notion-block-3a03936163f481c69611e6771a07b9a5" style="list-style-type:decimal"><li>给定当前自旋构型 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，对每条最近邻边采样 bond：若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，令 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则以概率 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 令 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-3a03936163f481ea8a50e8821cc69e71" style="list-style-type:decimal"><li>由所有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的边构造 FK clusters。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-3a03936163f48103b558dbe39966d4bc" style="list-style-type:decimal"><li>固定 bond 构型，对每个 cluster 独立采样一个自旋值 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，概率各为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并令 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对所有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 成立。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-3a03936163f481ed8e12cc318c3275ba" style="list-style-type:decimal"><li>重复以上两类条件采样，得到新的自旋构型序列。</li></ol><div class="notion-text notion-block-3a03936163f4814cad56c7deeb85054d">Wolff 算法是相同思想的单簇版本：随机选择一个种子自旋，只沿同向邻居以概率 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 生长一个 cluster，然后将该 cluster 整体翻转。与 Swendsen-Wang 相比，Wolff 每步只更新一个簇，但在二维临界 Ising 中通常同样十分高效。</div><div class="notion-text notion-block-39f3936163f48178ad1fede4c1bfa93a">Wolff 算法每步构造并翻转一个簇；Swendsen-Wang 算法则分解整个系统为多个簇并同时更新。由于 cluster 更新能够直接改变临界涨落中的大尺度结构，其动态临界指数通常显著小于局域更新，因此是二维 Ising 临界采样的标准高效方案。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f48102bc60f4da2ed728f0" data-id="39f3936163f48102bc60f4da2ed728f0"><span><div id="39f3936163f48102bc60f4da2ed728f0" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f48102bc60f4da2ed728f0" title="1.4 VAN 与神经网络采样（2019）"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1.4 VAN 与神经网络采样（2019）</span></span></h4><div class="notion-text notion-block-3a13936163f4819b9c10d7617cd21d3a">Variational Autoregressive Network（VAN）的主干不是把神经网络仅作为 Metropolis-Hastings 的 proposal，而是把自回归神经网络作为变分概率分布 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，直接近似目标 Boltzmann 分布 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。其优化目标是最小化反向 KL 散度</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a13936163f481b2a45ec4b750b41964">代入 Boltzmann 分布后，有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a13936163f481d388a3dd54342e7fe4">其中真实自由能为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而变分自由能为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a13936163f48145896df657846657df">由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，故 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。训练 VAN 即是降低 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，从而使 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 接近 Boltzmann 分布。</div><div class="notion-text notion-block-3a13936163f48116aac6c2735baa0186">自回归结构保证概率分布自动归一化。给定一个预先规定的格点顺序，VAN 将联合概率分解为条件概率乘积：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a13936163f4813aa69ce7ff81e368cf">因此，模型可以按顺序 ancestral sampling 生成构型，同时也能精确计算任意生成构型的归一化概率 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这使得 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可用来自模型自身的样本估计，而不需要预先准备来自目标 Boltzmann 分布的训练数据。</div><div class="notion-text notion-block-3a13936163f4817391b5ffabc7064288">参数更新可写成 score-function / policy-gradient 形式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3a13936163f48132a13de47acd7da712">实际实现中通常用 Monte Carlo 样本估计该梯度，并可加入 baseline 降低方差。因而 VAN 的核心流程是：用自回归网络定义归一化变分分布，直接从 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 采样，估计变分自由能及其梯度，然后优化网络参数。生成近似独立样本是训练后的结果之一，而不是该方法的定义性主干。</div><div class="notion-text notion-block-3a13936163f481369be9cd3b815d893d">在二维铁磁 Ising 这类 cluster 算法适用的模型中，Wolff 和 Swendsen-Wang 通常仍是更标准、更稳健的临界采样基准。VAN 更适合作为变分求解与生成式采样框架，尤其适用于传统 cluster 构造困难、分布多峰或存在复杂约束的模型；其代价是训练过程本身可能成为主要计算瓶颈。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39f3936163f481ecab48fc6288a6000d" data-id="39f3936163f481ecab48fc6288a6000d"><span><div id="39f3936163f481ecab48fc6288a6000d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481ecab48fc6288a6000d" title="2. Kadanoff 数值重整化群的实施方案"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. Kadanoff 数值重整化群的实施方案</span></span></h3><div class="notion-text notion-block-39f3936163f481a8b501eef1f6fa3e8d">Kadanoff 数值重整化群的基本任务，是将微观格点模型的 Boltzmann 分布通过给定的粗粒化映射推送到粗变量空间，并在适当的算符基底中构造相应的有效哈密顿量。若需要显式求出粗粒化后的耦合常数，标准数值框架是 Monte Carlo Renormalization Group，常用实现为 inverse Monte Carlo 或 Swendsen 方法。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481feb3cbeb02bded853b" data-id="39f3936163f481feb3cbeb02bded853b"><span><div id="39f3936163f481feb3cbeb02bded853b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481feb3cbeb02bded853b" title="1. 微观模型与算符表示"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. 微观模型与算符表示</span></span></h4><div class="notion-text notion-block-39f3936163f4815aa365dfcfeaa2cbb4">设微观构型记为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在一组算符 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上，无量纲哈密顿量可写为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481caba3ff5ee71cce77b">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 为耦合常数，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可取最近邻能量、次近邻能量、磁场项或多体相互作用项。以下符号约定均采用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；若采用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，线性响应公式中的符号需相应调整。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f4810fa15fc75d8d27f02a" data-id="39f3936163f4810fa15fc75d8d27f02a"><span><div id="39f3936163f4810fa15fc75d8d27f02a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f4810fa15fc75d8d27f02a" title="2. 微观系综的数值采样"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. 微观系综的数值采样</span></span></h4><div class="notion-text notion-block-39f3936163f481df880fcb8b8488c1ce">在给定裸耦合 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 下，通过 Monte Carlo 方法产生满足 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的构型序列 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在临界点附近，应优先采用 Wolff 或 Swendsen-Wang 等 cluster update，以减小临界慢化的影响。采样前需要充分热化，采样后需要估计自相关时间，并据此给出统计误差。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f48194850ac0931e6e9b2a" data-id="39f3936163f48194850ac0931e6e9b2a"><span><div id="39f3936163f48194850ac0931e6e9b2a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f48194850ac0931e6e9b2a" title="3. 粗粒化映射"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 粗粒化映射</span></span></h4><div class="notion-text notion-block-39f3936163f481c0b90cddd17edb2bcb">选定尺度因子 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 以及从微观变量到粗变量的映射 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，即</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481f2a7d5f1b6d28b1cda">例如二维 Ising 模型中可采用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> majority rule。若块内正负自旋数相等，需要事先规定无偏的 tie-breaking 规则。重复实施 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 次粗粒化后，线性尺度变为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f48160a4dfc9e3b1291597" data-id="39f3936163f48160a4dfc9e3b1291597"><span><div id="39f3936163f48160a4dfc9e3b1291597" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f48160a4dfc9e3b1291597" title="4. 粗粒化系综"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. 粗粒化系综</span></span></h4><div class="notion-text notion-block-39f3936163f48128ab71da58ef4b674e">对每个微观样本施加相同的粗粒化映射，得到粗格点构型序列</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481cab1b7ddf301dd1153">从概率分布的角度看，RG 变换定义为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f48123820dcba44bdbb7a7">由此得到的粗格点构型集合称为 blocked ensemble。粗格点算符的样本平均均在该系综上计算。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481ecbb45f86b549ebc6b" data-id="39f3936163f481ecbb45f86b549ebc6b"><span><div id="39f3936163f481ecbb45f86b549ebc6b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481ecbb45f86b549ebc6b" title="5. 精确定义的有效哈密顿量"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. 精确定义的有效哈密顿量</span></span></h4><div class="notion-text notion-block-39f3936163f481668ad2dcf1540f10c9">粗粒化后的有效哈密顿量由以下恒等式定义：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f48113b4c6ce7adc4cebaf">该表达式给出了 Wilson-Kadanoff 意义下的精确有效作用量。然而，对一般格点系统，上式所含的微观构型求和无法直接完成，因此需要在数值上引入算符截断。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f4818c8c74e902cf0f4dbf" data-id="39f3936163f4818c8c74e902cf0f4dbf"><span><div id="39f3936163f4818c8c74e902cf0f4dbf" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f4818c8c74e902cf0f4dbf" title="6. 有效作用量的算符截断"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">6. 有效作用量的算符截断</span></span></h4><div class="notion-text notion-block-39f3936163f481c184d0dfee7c3341bf">在粗格点上选取一组有限的局域算符，并令</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f48166b996e8f2ecc6d8cf">即使原始模型只含最近邻相互作用，粗粒化后也会生成所有由对称性允许的项，包括次近邻项、多体项与更长程项。以 Ising 型变量为例，可写为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481aa825af2a7822f2e8b">因此，数值 RG 中得到的耦合流通常是精确 RG 流在有限算符子空间中的投影。算符基底的选择直接决定截断误差。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f48178851ef367c7190b79" data-id="39f3936163f48178851ef367c7190b79"><span><div id="39f3936163f48178851ef367c7190b79" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f48178851ef367c7190b79" title="7. 耦合常数的反演"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">7. 耦合常数的反演</span></span></h4><div class="notion-text notion-block-39f3936163f481d8a3b1d5e2c2837b5e">inverse Monte Carlo 的基本条件是矩匹配：有效模型中的算符平均应等于 blocked ensemble 中的样本平均，即</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f4819cac85c1c769753d91">定义</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f4815cad39ed981a73ac45">根据线性响应关系，有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f48120846bcc8da7f3ab4b">因此，Newton 迭代给出 Swendsen inverse Monte Carlo 公式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481f7b404d1acf61b87a0">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 为候选有效模型中的算符协方差矩阵。它同时给出各算符的涨落尺度与相互相关性，因此其逆矩阵在迭代中起到预条件化作用。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481c1a4b2eb3dd827eb34" data-id="39f3936163f481c1a4b2eb3dd827eb34"><span><div id="39f3936163f481c1a4b2eb3dd827eb34" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481c1a4b2eb3dd827eb34" title="8. RG 流与临界指数"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">8. RG 流与临界指数</span></span></h4><div class="notion-text notion-block-39f3936163f48196b240e5400a17d296">对同一微观系综连续实施粗粒化，并在每一级上反演有效耦合，可得到耦合空间中的 RG 流：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f481d29657f5ae14bdbf75">在固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近，线性化 RG 变换为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f48137ae32dfad2c1650ea">若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的本征值，则相应 RG 指数为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39f3936163f4811e8f45e66c919cb109">热方向指数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 与相关长度指数满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481449600edf4602b9490" data-id="39f3936163f481449600edf4602b9490"><span><div id="39f3936163f481449600edf4602b9490" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481449600edf4602b9490" title="9. 二维 Ising 模型中的最小实现"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">9. 二维 Ising 模型中的最小实现</span></span></h4><ol start="1" class="notion-list notion-list-numbered notion-block-39f3936163f481678452e4322062f1fa" style="list-style-type:decimal"><li>在临界点附近采样二维方格 Ising 模型，优先采用 Wolff 或 Swendsen-Wang 更新。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-39f3936163f48155bafaeb32db912a3a" style="list-style-type:decimal"><li>经过热化后保存近似独立的构型样本，并估计自相关时间。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-39f3936163f48165b087f5fb01d49855" style="list-style-type:decimal"><li>对每个构型施加 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> majority-rule blocking，并可重复多级粗粒化。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-39f3936163f481e3984dfc131fd56b58" style="list-style-type:decimal"><li>在粗格点上测量最近邻、次近邻、plaquette 四体项、磁化等算符。</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-39f3936163f481479cc2ccca54f932d5" style="list-style-type:decimal"><li>选定有限算符基底，并通过 inverse Monte Carlo 解矩匹配方程，得到各级 blocking 后的有效耦合。</li></ol><ol start="6" class="notion-list notion-list-numbered notion-block-39f3936163f48199b9b7e03fa45b1b99" style="list-style-type:decimal"><li>比较不同系统尺寸与不同 blocking level 的结果，以检验有限尺寸效应和固定点附近的稳定性。</li></ol><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f481658daacf66b98ec226" data-id="39f3936163f481658daacf66b98ec226"><span><div id="39f3936163f481658daacf66b98ec226" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f481658daacf66b98ec226" title="10. 误差来源与一致性检验"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">10. 误差来源与一致性检验</span></span></h4><ul class="notion-list notion-list-disc notion-block-39f3936163f4811796d5d52280112c9e"><li>算符截断误差：有限算符基底只能给出精确 RG 流在该子空间中的投影。</li></ul><ul class="notion-list notion-list-disc notion-block-39f3936163f481a896cff82010be31b4"><li>有限尺寸误差：过多的 blocking 会使有效格点数过小，从而破坏热力学极限近似。</li></ul><ul class="notion-list notion-list-disc notion-block-39f3936163f48150a931dead1bf94df7"><li>协方差矩阵病态：算符强相关会导致 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 放大统计噪声，可通过删去冗余算符、增加样本量或引入正则化缓解。</li></ul><ul class="notion-list notion-list-disc notion-block-39f3936163f481cda9cff900985e7f9b"><li>符号约定误差：不同哈密顿量记号会改变响应方程中的正负号。</li></ul><ul class="notion-list notion-list-disc notion-block-39f3936163f481f3b614d9c546041266"><li>粗粒化规则依赖：不同 blocking scheme 会改变非普适耦合，但临界指数等普适量应保持一致。</li></ul><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39f3936163f48186862bc2b3d218360c" data-id="39f3936163f48186862bc2b3d218360c"><span><div id="39f3936163f48186862bc2b3d218360c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39f3936163f48186862bc2b3d218360c" title="结论"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">结论</span></span></h4><div class="notion-text notion-block-39f3936163f4813ea016cac3b5f3fffc">Kadanoff 数值重整化群可概括为：采样微观系综，施加粗粒化映射得到 blocked ensemble，在有限算符基底中用矩匹配或 inverse Monte Carlo 反演有效耦合，并由连续 blocking 得到 RG 流、固定点及临界指数。</div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Renormalization Group]]></title>
            <link>https://blog.xiangsiqi.site/RenormalizationGroup</link>
            <guid>https://blog.xiangsiqi.site/RenormalizationGroup</guid>
            <pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39e3936163f480d6a7d3d58517284395"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-39e3936163f481edb720e2d1d5770ee6" data-id="39e3936163f481edb720e2d1d5770ee6"><span><div id="39e3936163f481edb720e2d1d5770ee6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481edb720e2d1d5770ee6" title="1. Problems"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Problems</span></span></h2><div class="notion-text notion-block-39e3936163f481a1aa4ae05a68ac5b53">Renormalization Group 的出发点不是先引入一套新的计算技巧，而是先把 Landau-Ginzburg 理论中真正没有解决的问题说清楚。Landau-Ginzburg 已经把序参量从一个均匀变量推广成空间场，因此它能描述关联函数和关联长度；可是如果我们仍然只在高斯近似或鞍点附近工作，临界指数仍然停留在平均场结果。问题是：临界点附近的涨落为什么不能被当作小修正？什么时候平均场会失效？</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39e3936163f4810ea68ce05fb9aa13a5" data-id="39e3936163f4810ea68ce05fb9aa13a5"><span><div id="39e3936163f4810ea68ce05fb9aa13a5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f4810ea68ce05fb9aa13a5" title="1.1 Landau-Ginzburg ：平均场与它的失败"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1.1 Landau-Ginzburg <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：平均场与它的失败</span></span></h3><div class="notion-text notion-block-39e3936163f481f8a1c8d284eb21cc7d">对 Ising 型二级相变，Landau-Ginzburg 理论把序参量写成一个实标量场。最简单的自由能泛函是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b78c4ae5307df6705d">这里的梯度项惩罚空间不均匀，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 控制距离临界点的远近，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 保证自由能在大场强下稳定，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是与序参量共轭的外场。若忽略空间涨落，只取均匀场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，就回到 Landau 平均场理论：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481958e55efa16b26fc25">在零外场下，极值条件给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4816684c8ea38ba3b0ac2">因此高温侧 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时稳定解是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；低温侧 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时出现两个等价极小值</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b4b3e5fd1341d43c33">这给出平均场指数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。类似地，高温侧的均匀磁化率来自</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481588042e3c47d6849bb">所以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；在临界等温线 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。若进一步保留高斯涨落，在高温侧忽略 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 项，可得</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4817d8789e2b5e5827e89">于是动量空间关联函数为 Ornstein-Zernike 形式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f188e3ee15fa9a44b5">这说明 Landau-Ginzburg 理论比均匀 Landau 理论前进了一步：它已经看见了空间关联长度，并给出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。然而这些指数仍然是平均场指数。原因是高斯理论只描述彼此独立的 Fourier 模式，而真正的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 相互作用会让不同波长的涨落耦合起来。</div><div class="notion-text notion-block-39e3936163f481f4b617cd962627a124">平均场失败的物理原因可以这样说：它假设临界点附近的主要行为由一个代表性平均值控制，涨落只是围绕这个平均值的小扰动。但临界点附近 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 发散，一个涨落团簇可以跨越任意大的尺度；此时没有一个固定的局域尺度可以把涨落平均掉。系统不是许多近似独立的小区域，而是多尺度涨落彼此嵌套的整体。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39e3936163f481128e40ffc0e4c7bfee" data-id="39e3936163f481128e40ffc0e4c7bfee"><span><div id="39e3936163f481128e40ffc0e4c7bfee" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481128e40ffc0e4c7bfee" title="1.2 Ginzburg criterion：什么时候涨落不可忽略"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1.2 Ginzburg criterion：什么时候涨落不可忽略</span></span></h3><div class="notion-text notion-block-39e3936163f4812b8060c7ea9c75db91">Ginzburg criterion 的目的，是把“涨落很重要”这句话变成一个可检验的不等式。思路是比较相关体积 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 内的平均场序参量平方，和同一体积内的热涨落强度。如果涨落远小于平均场值，平均场自洽；如果涨落和平均场值同阶，平均场就不能作为起点。</div><div class="notion-text notion-block-39e3936163f481a69426c80417294a1a">先看低温侧。平均场给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481849614de2d5c341a88">一个相关体积内的粗粒化序参量可以写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d9aa67d360bdcf886a">它的涨落量级由关联函数积分估计：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4812e9705f9f56310b56c">这一步可以展开得更清楚。先定义局域涨落 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则粗粒化平均场的涨落是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d59193f06c2dd02c55">所以方差是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4811da5c0e652e74d785f">连通关联函数定义为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ad9a64f38094c3d504">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48123900bfbce74a66579">现在利用平移不变性。对固定的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，令相对坐标 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。如果忽略相关体积边界附近的非主导修正，内层积分的量级为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481e185c0fbeb424b6991">外层 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 积分只给出一个体积因子 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，于是双重积分不是给出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而是给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48105a961c37e93132601">代回方差表达式，就得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d3aac7f90c0a560377">最后，磁化率本身就是连通关联函数的空间积分，差别只在于是否把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 这样的常数吸收到定义中：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481709797dcd5c199da7d">因为离临界点但接近临界点时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 后被指数衰减截断，所以积分的主要贡献来自 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此，也可以把相对坐标重新记为通常的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48140b852cf86c4e734b1">所以相关体积平均后的涨落量级就是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b0a48fe3c1104310be">在高斯近似中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481fe8481d4d4a4581d59">平均场要求这个涨落远小于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48118a8b7e7732e97610e">代入上面的量级估计，得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f69da1fa5add148022">这就是 Ginzburg criterion 的核心形式。它告诉我们：当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 反而使左边趋于 0，平均场在临界点附近越来越可靠；当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 会使左边发散，说明无论 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 多小，只要足够接近临界点，涨落最终都会压倒平均场。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ba96bfcf4e359e6eac">因此四维是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 理论的上临界维数。高于四维，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 相互作用在长距离下变得不重要，Gaussian fixed point 控制临界行为，平均场指数成立；低于四维，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 相互作用在长距离下不可忽略，必须寻找超出 Gaussian fixed point 的新临界理论。</div><div class="notion-text notion-block-39e3936163f481eca77dcb9eb4c3aa59">这正是 Renormalization Group 要解决的问题：不是一次性计算所有涨落，而是按长度尺度逐层处理涨落，观察参数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 等在粗粒化下如何流动。Kadanoff block spin 给出这个思想的物理图像，Wilson RG 把它变成可计算的动量壳积分，而 Wilson-Fisher fixed point 则是在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时替代平均场的非平庸临界固定点。</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-39e3936163f4811c9c62f12295eb24f0" data-id="39e3936163f4811c9c62f12295eb24f0"><span><div id="39e3936163f4811c9c62f12295eb24f0" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f4811c9c62f12295eb24f0" title="2. Kadanoff to Wilson"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. Kadanoff to Wilson</span></span></h2><div class="notion-text notion-block-39e3936163f481cb82cec09fb11c5d25">第一节说明了为什么平均场不够：临界点附近的涨落不是局域小修正，而是跨越许多长度尺度的集体现象。第二节的任务，是把“多尺度涨落”变成一种系统方法。Kadanoff 给出实空间中的粗粒化图像；Wilson 则把这个图像改写成连续场论中可计算的动量壳积分。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39e3936163f481a7a46edce99a3d2b29" data-id="39e3936163f481a7a46edce99a3d2b29"><span><div id="39e3936163f481a7a46edce99a3d2b29" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481a7a46edce99a3d2b29" title="2.1 Coarse Instantiation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2.1 Coarse Instantiation</span></span></h3><div class="notion-text notion-block-39e3936163f48126b49ed38df9bdb191">Landau-Ginzburg 模型本身就已经是粗粒化后的有效理论。我们不再直接操作微观自旋，而是先从微观模型 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 出发，把一小块区域里的自旋平均成连续序参量场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481478c96f7feb0f19dcd">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是粗粒化长度，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 为中心、线度为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的小盒子。于是我们把所有满足同一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的微观构型加总掉，得到这个粗粒化场的有效自由能：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f98b8de6448b4e9cf7">所以从这一刻起，后面的操作对象都不是原始微观哈密顿量，而是这个有效自由能 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。Landau-Ginzburg 泛函只是对它再做长波展开后的标准形式。换句话说，RG 不是先从裸微观模型直接开始，而是先站在这个 coarse-grained effective theory 上，再讨论 Kadanoff block spin 和 Wilson coarse-graining。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39e3936163f481a8aee3ec1ed2eb8efe" data-id="39e3936163f481a8aee3ec1ed2eb8efe"><span><div id="39e3936163f481a8aee3ec1ed2eb8efe" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481a8aee3ec1ed2eb8efe" title="2.2 Kadanoff block spin"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2.2 Kadanoff block spin</span></span></h3><div class="notion-text notion-block-39e3936163f4812e957bfe850afd0d90">Kadanoff 的基本想法是：临界点附近关联长度 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 很大，系统中出现许多大小不同的相关团簇。因此我们不应该只盯着单个格点自旋，而应该把一小块区域看成一个新的有效自由度。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f4811bb428d1b8e2e4d897" data-id="39e3936163f4811bb428d1b8e2e4d897"><span><div id="39e3936163f4811bb428d1b8e2e4d897" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f4811bb428d1b8e2e4d897" title="Block average calculation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Block average calculation</span></span></h4><div class="notion-text notion-block-39e3936163f48100937bd4c2590c822e">把 block spin 的定义先换成连续场语言。取线度为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的小盒子 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，定义块平均场</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f39994daecabba27e7">这是 Kadanoff coarse-graining 的核心：局域变量不是由单点决定，而是由一个体积平均决定。现在把它做 Fourier 变换，就能看到短波模式如何被压掉。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4816bae05eeb55730665c">这个公式来自矩形窗口函数的 Fourier 变换。若把盒子写成以 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 为中心，令 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481658033fdc07b426f14">把原场展开为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481309e67defd8e636fff">代回块平均定义：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4810f93afe5ef143ffdd4">括号中的量就是窗口函数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48104b74bf5df1bbb1fce">单个方向的积分为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481e491a3d100426732ec">因此对于居中盒子，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48113af5eeef50e3a5471">如果盒子取为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 而不是居中盒子，则每个方向多出一个整体相位 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或其复共轭，取决于 Fourier 约定。这个相位不影响 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对高动量模式的抑制；真正重要的是 sinc 因子。</div><div class="notion-text notion-block-39e3936163f4818d9111d43e9acc47cd">这个 form factor 在长波极限下接近 1，在短波极限下强烈抑制高动量模式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4812f9f9dcdc5da3d3ed9">这里</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48166ad32e14a849f2e56">所以 block averaging 自动给出一个有效 cutoff <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：高于这个尺度的 Fourier 分量被平均掉，留下的只剩慢变的长波模式。</div><div class="notion-text notion-block-39e3936163f481e59b89dcc827b81da6">上面的 form factor 说明了 block averaging 在动量空间中如何压低短波模式。接下来从实空间再看同一件事：block average 后的变量涨落幅度是否真的变小？这就需要计算块平均场的方差。若方差随 block 体积增大而下降，说明粗粒化后的变量确实更平滑；若临界点附近方差不按普通中心极限定律下降，就说明长程相关已经破坏了“块内自由度近似独立”的直觉。</div><div class="notion-text notion-block-39e3936163f4815aade4ef32b3b7da9f">从定义出发，块平均场为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ee88f5ddddd6cdb20d">定义局域涨落和块平均涨落：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481bdb039e4b5929cf4c8">因为平均操作是线性的，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48160bd86d0c904ea17e3">于是块平均场的方差为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4819e85fac7e4b6b0ef31">展开平方，把平均值移入积分：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48134bf3debc8af85cd64">再用连通关联函数</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48169b51cc26a9fab44c2">就得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b1866acf994b217dae">若远离临界点，相关长度远小于 block 尺度，可以把关联函数近似成局域相关：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b68153eb2084b9f348">代入方差公式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481358c1acaa0bd6e3e6d">先对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 积分：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4818584c9fd7f0fab6f20">于是只剩一个体积积分：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ff993ccc512bc11013">这就是中心极限定律的连续场版本：一个 block 内大约有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个近似独立自由度，平均后的方差按 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 缩小。</div><div class="notion-text notion-block-39e3936163f4813a8d00c2e336d5a5bc">更一般地，只要关联长度有限，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，就不必真的把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 写成严格的 δ 函数。由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 只在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 内显著，对大 block 可近似为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48133a29ac77e1a86ddae">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a9840aed731405de1a">这说明真正控制方差缩放的不是单点噪声本身，而是关联函数在一个 block 内的积分。如果到临界点，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 发散，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在块内不能看成短程函数。例如临界点附近</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48100a5a9def1fee6264e">这时</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481be8e1df334e2717f61">这说明 block 平均确实让变量更平滑、更少涨落；但在临界点附近，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，相关函数不再是 δ 函数，方差不再简单按 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 缩小，于是多尺度相关性就必须进入 RG。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f48117b84fda2ba465fdf6" data-id="39e3936163f48117b84fda2ba465fdf6"><span><div id="39e3936163f48117b84fda2ba465fdf6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f48117b84fda2ba465fdf6" title="Generation of operators in the Ising model"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Generation of operators in the Ising model</span></span></h4><div class="notion-text notion-block-39e3936163f48159a4bfc6253ed609db">Kadanoff coarse-graining 的最直观形式，是先把原格点分成互不重叠的 blocks <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，每个 block 先取平均磁化，再据其符号定义新自旋。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481baa773c7095908ed3c">若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，等价于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，可以约定随机取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，或采用固定 tie-breaking rule。这样得到一个从微观构型到粗粒化构型的映射</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4812e9e9ef76107b97361">现在定义粗粒化后的有效 Hamiltonian。固定一个 coarse configuration <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，把所有会被 block map 映到它的微观构型求和：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48161b32adf3af430f26c">这就是“对 block 内部自由度求和”的精确定义：不是对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 直接平均，而是对 Boltzmann weight 求和，然后用负对数定义新的有效 Hamiltonian。</div><div class="notion-text notion-block-39e3936163f481e685add8cccbd4475c">同一个 constrained sum 也可以写成 delta 约束形式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4818ab4b4c9de4117b18a">将这一硬约束进一步写成 block-spin kernel，便得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4817a8ee6ec144b982de7">先看为什么一般 RG 流必须写在无穷多耦合的空间里。以最近邻 Ising 模型为起点：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814cb830c498a147fa35">在有限体积中，任意函数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 都可以严格展开在 Ising 变量的乘积基上。若 coarse-grained lattice 有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个格点，则</div><div class="notion-text notion-block-39e3936163f4812d9939f7661e1dd41c">这句话的含义是一个有限维函数空间的代数事实。若粗粒化后的格点数为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是定义在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个 Ising 构型上的函数：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c784b3d44a31ab6bf5">因此这个函数空间的维数就是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。记粗粒化后的格点集合为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，其中每个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 标记一个 coarse-grained spin <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。对任意子集 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，即任意一组粗粒化格点，定义乘积函数</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48161bdb6c257f3b7f4b3">例如 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。子集 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的数目也是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并且这些乘积函数彼此正交。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4819c97f6d5bf1afafdaa">这是因为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 构成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上函数空间的一组正交基：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4817185a3e4caa2e255f3">证明如下。两个乘积函数相乘时，重复出现的自旋平方为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以只剩下属于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 或属于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 但不同时属于二者的自旋：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c4a934f3f7f114aac2">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是 symmetric difference，即只属于其中一个集合的格点。若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，乘积恒等于 1，因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481df95c7c27745ea7899">若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。取一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，求和中含有因子 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。对所有 Ising 构型求和时，固定其他自旋不变，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的两项正负抵消，所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ae8ff6e86370b207bd">合并两种情况，得到正交归一关系</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48137bba8e0fd6884a276">因此系数由投影公式给出：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481428250fc6787fe4a9c">这已经说明：即使原始 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 只含最近邻二体项，粗粒化以后也会生成更高阶的耦合。</div><div class="notion-text notion-block-39e3936163f48162b128f1c38ae44037">下面为了把这一点具体算出来，取二维 square-lattice Ising 的最简单 decimation 作为例子。Decimation 并不是标准的 Kadanoff block-spin majority rule：它不是把一个 block 平均或投影成一个新自旋，而是保留一个子格点上的原自旋作为新的 coarse variables，把其余子格点求和掉。这里采用 decimation，是因为它能精确展示一次粗粒化如何生成新的有效相互作用。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39e3936163f4814c851ff63e26a4cb97"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://www.notion.so/image/attachment%3Acc98b036-a5da-4830-91a8-b4391bc0ced6%3Aising_decimation_schematic.png?table=block&amp;id=39e39361-63f4-814c-851f-f63e26a4cb97&amp;t=39e39361-63f4-814c-851f-f63e26a4cb97" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39e3936163f481c9879fe7f2b16304d7">对图中的局部星形簇，消去中心黑格点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 后，四个保留自旋 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的局部有效权重定义为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4819dbdebf69fa5efc5a4">原来的最近邻 Hamiltonian 在这个局部簇上的贡献为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481e79e68fb9116d96aa0">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a490c3e4fbbc4f1a41">这个结果已经只依赖保留自旋，下一步是把它重新写成这些保留自旋之间的有效相互作用。由对称性，最一般的局部形式可以写为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a4ab4dee011c34bf19">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是写在 Boltzmann weight 指数中的有效耦合；若写回 Hamiltonian，则对应局部贡献 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-39e3936163f481889ae6ec559e1aac3d">所有系数都可以直接用 Ising 乘积基的正交性求出。令</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481fd95b2ee2d27f214c7">由于指数右边就是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在乘积基上的展开，各个系数由正交投影给出：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48134a0d9da4e20b8b1b3">这里每个求和都遍历 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的全部 16 个构型。记</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481cfa2c7dc0a406d2707">按照 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 分类，构型数分别为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此常数项为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c19ef9cec4f13102a7">对边耦合，投影到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上。直接枚举给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d1acaadeb4dbf0b0eb">对角耦合由正方形对称性或直接投影到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 得到相同结果：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48164be20d4620fe6bfe9">四自旋耦合投影到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b98316dedfdefe7096">所以这一局部 decimation 的完整系数为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a58b87e006c8ef1830">所以一次精确 decimation 已经从最近邻 Ising 模型生成了保留格点之间的多种二体耦合和四自旋耦合。继续迭代时，不同局部因子相互重叠，还会生成更远程和更高阶的项。因此后面的记号</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c49298e49b7ccc65d7">不是形式上的装饰，而是 Ising 模型在粗粒化下不封闭的必然结果。这里的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示 decimation 后保留下来的 coarse variables，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是这些保留自旋的乘积项，例如二体项、四体项、更远程项等。进入下一小节讨论一般 RG flow 时，可以把当前层变量重新记作 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，于是写成通用形式 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39e3936163f481c3b764d801db532466" data-id="39e3936163f481c3b764d801db532466"><span><div id="39e3936163f481c3b764d801db532466" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481c3b764d801db532466" title="2.3 RG flow, Fixed Point and Critical Point"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2.3 RG flow, Fixed Point and Critical Point</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f48115b1d6d8c5f794ce3a" data-id="39e3936163f48115b1d6d8c5f794ce3a"><span><div id="39e3936163f48115b1d6d8c5f794ce3a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f48115b1d6d8c5f794ce3a" title="RG flow"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">RG flow</span></span></h4><div class="notion-text notion-block-39e3936163f481419a6bd939eb05e9ab">严格地说，本小节从这里开始已经不再只是 Kadanoff 原始意义上的 block-spin picture。Kadanoff 的核心是实空间 coarse-graining：把局域自由度组合成更粗的变量，并说明临界附近应当出现尺度变换。下面引入耦合空间 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>、RG 映射 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>、固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>、本征方向和临界指数时，实际上是在使用 Wilson 对 Kadanoff 思想的系统化：real-space Wilson RG。</div><div class="notion-text notion-block-39e3936163f4812aaf71c42ba0d50ef6">因此本节的逻辑不是把所有内容都归为“纯 Kadanoff 计算”，而是先用 Kadanoff block spin 建立粗粒化的物理图像，再把它提升为 Wilson 式的 RG flow。下一节的 momentum-shell RG 是同一 Wilson 思想在动量空间中的可计算实现。</div><div class="notion-text notion-block-39e3936163f4815e8eacc4fd4ef70cf8">为了把 Kadanoff block spin 写成真正的 RG flow，不应只保留最近邻 Ising 耦合。设当前 coarse-grained Hamiltonian 已经展开在一组允许的算符上：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481bcaf5bf58bd801c976">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是耦合常数空间；<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 包括最近邻、次近邻、多自旋项、长程项等所有在 coarse-graining 下会生成的相互作用。block-spin 规则可用一个条件概率核表示：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4813a97c5f90deb6ffcb1">例如 majority rule 对应把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 固定为 block <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 内自旋和的符号；更一般地，也可以采用随机化的 coarse-graining 规则，使 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 成为非平凡的条件概率分布。</div><div class="notion-text notion-block-39e3936163f4816fa4c1d9df0ef984e6">给定这样的 block-spin kernel，做一次尺度因子为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的粗粒化后，变量从 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 变成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，Hamiltonian 重新写成同一类算符展开，但耦合常数从 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 变为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4816d977bcffbf45565f7">因此 RG flow 主要追踪的不是某一个具体自旋构型如何运动，而是耦合空间中的点如何变换：<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。下面的 constrained trace 正是这个映射的定义。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814c8684d2895be8760c">上式定义了耦合空间中的映射</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4812fb9f3e6643e32ba70">严格地说，单纯 coarse-graining 会把格距从 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 变成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此先得到的是定义在更粗格子上的 Hamiltonian。只有再做长度重标度，把新的格距重新作为单位格距，并把粗粒化后的变量重新命名为当前层变量，才可以把新旧理论看成同一个耦合空间 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中的两个点。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814daed5f422083101e7">因此写成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，已经包含了 coarse-graining 之后的长度重标度这一步。</div><div class="notion-text notion-block-39e3936163f4813fa288ea5fcf0f54c1">由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，配分函数保持不变，差别至多是 Hamiltonian 中一个与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 无关的加性常数：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48113bf50edf35e306035">这个等式说明 RG 不是丢弃热力学，而是把同一个配分函数重新写成更大晶格间距 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上的有效理论。为了理解 RG flow 如何区分临界点和非临界点，进一步考察相关长度在一次 RG 变换下如何变化。</div><div class="notion-text notion-block-39e3936163f481f1b3b5fe0cabb773d4">从定义出发。设当前格距为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，连通关联函数在远距离的指数衰减定义了物理相关长度：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c49aaac1aa43d6d0d3">若用格距作为长度单位，记无量纲相关长度为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ce9d60cb16a430ef98">一次尺度因子为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 coarse-graining 把格距变为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。新的关联函数由粗粒化变量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 定义；但只要 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是原变量的局域粗粒化，并且与原序参量有非零 overlap，它的长距离连通关联函数具有同一个物理指数衰减长度。coarse-graining 改变短距离振幅和局域细节，长度重标度改变的是度量这个长度的格距单位：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481379c30c77df773b6a8">于是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4812b8df6e4d0eef0d105">因此非临界点的 RG 流是严格不同的。若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，反复粗粒化 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 次后</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f481a8b2e2c2a2b7d03b94" data-id="39e3936163f481a8b2e2c2a2b7d03b94"><span><div id="39e3936163f481a8b2e2c2a2b7d03b94" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481a8b2e2c2a2b7d03b94" title="Fixed Point"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Fixed Point</span></span></h4><div class="notion-text notion-block-39e3936163f481cdaaaac5e843cf2f98">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，重标度后的系统只剩 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 格点尺度的相关性。高温相流向无序固定点，低温相流向有序固定点：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481e2a23ad724be572927">高温固定点可以更具体地理解。这里的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是无量纲耦合常数；对最近邻 Ising，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，所以高温极限对应 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。若所有相互作用耦合都为零，</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f9aa95f9edc885d38a">则所有微观构型的 Boltzmann weight 相同。代入 coarse-graining 定义，得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481efb44ee6ac628cd601">对于平移对称的 coarse-graining rule，右边只给出与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 无关的计数因子，因此 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 仍只是常数。也就是说，无相互作用点是 RG 固定点：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a0905ed15b26fcaff8">高阶耦合虽然会在粗粒化中生成，但在高温固定点附近它们是原始小耦合的高阶函数，因此随 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 一起消失。前面的 decimation 例子给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b4868efa3f4133e32b">利用小量展开</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481de90e7e256c2c2bf3a">得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481daa920c07b35d8a3bc">因此在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 处所有非平凡相互作用耦合严格为零；在高温相内，反复 RG 会把有限相关长度的系统推向这个无序固定点。</div><div class="notion-text notion-block-39e3936163f481399f90c8dc97413162">低温固定点的行为不同。低温极限对应 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，此时相互作用不是趋于零，而是进入强耦合区域。仍以前面的 decimation 系数为例，利用</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ab94d2e91c90cc7986">得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48160aa78cbf390a8c5ea">所以低温固定点不是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而是强耦合固定点。更清楚的变量是 domain-wall fugacity。局部权重为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4811c9d5fdc733e8a2494">当四个保留自旋全同号、三同一反、两正两负时，权重分别满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ed9b32c38f29d07aa8">相对于全同号构型，含 domain wall 的构型被指数压低：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c18dcff51410addb83">因此低温相的自然小参数不是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。低温 RG 流可理解为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，也就是 domain wall 的权重消失，系统流向有序固定点。</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f48123bc96c3e18802961f" data-id="39e3936163f48123bc96c3e18802961f"><span><div id="39e3936163f48123bc96c3e18802961f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f48123bc96c3e18802961f" title="Critical Exponent"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Critical Exponent</span></span></h4><div class="notion-text notion-block-39e3936163f48198be5ff83add7df2ee">临界点的特殊性在于</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48154a698e69956b92366">所以临界集合</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481428f5ed1b2feecfdc8">在 RG 下不流向平庸有限相关长度理论，而是在临界面内流动。若存在非平庸固定点，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c6ab96eb3002a3bb17">这才是 Kadanoff 图像中尺度不变性的数学内容：在固定点处，coarse-graining 加上长度重标度后理论回到自身。</div><div class="notion-text notion-block-39e3936163f481a9be56ca40ef35a771">接下来要研究的不是固定点本身，而是临界固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近的流。这个稳定性分析回答三个问题：哪些相对于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的扰动会把系统带离临界面，哪些扰动会在长距离下被遗忘，以及相关长度、磁化率等临界指数由哪些尺度因子决定。因此 relevant/irrelevant directions 和临界指数都来自固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近 RG 映射的线性化。</div><div class="notion-text notion-block-39e3936163f48133a3f0ebeec812a8e6">临界指数来自固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近的线性化。令</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481dca8a5e2de13feca5f">最低阶为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481dca835d348b7f2ce03">取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的本征方向作为固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近的坐标方向。也就是说，把偏离量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 展开到这些本征方向上，第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个本征方向上的增量坐标记为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。在这组坐标中，RG 线性阶不再混合不同方向，而是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814788a0fce988c6af95">若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则本征坐标 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在粗粒化下按 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 放大，称为 relevant；若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，该本征坐标被压低，称为 irrelevant；若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，需要看高阶项，称为 marginal。因此，要让一个点在 RG 下留在临界固定点的吸引流形上，所有 relevant 本征坐标必须调为零；否则任意非零 relevant 分量都会被放大，使流离开临界面并最终进入高温或低温平庸固定点。irrelevant 本征坐标则不需要调零，因为它们会在迭代中自动衰减。于是临界面在固定点附近由所有 relevant 本征坐标的调零条件给出：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d6aee5d204b12b8084">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是临界固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 附近的局部临界面，也就是该固定点的稳定流形：沿相对于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 irrelevant directions 可以有非零偏离，但这些偏离会被 RG 消去；沿相对于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 relevant directions 则必须精确调零。</div><div class="notion-text notion-block-39e3936163f481c7b08ffd086aa38ade">对于具有 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 对称且外场为零的 Ising 临界点，通常只有 thermal scaling field <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 需要调零；若允许外场，则 magnetic scaling field <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 也是 relevant。</div><div class="notion-text notion-block-39e3936163f4813aa3c0d7a4221fe3e9">普适性现在可以精确表述为：不同微观模型对应 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中不同初始点；只要它们的 RG 轨道流向同一个固定点 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并且具有相同的 relevant directions，所有 irrelevant 坐标都会按 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 衰减，所以长距离临界行为只依赖固定点数据 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而不依赖微观耦合的细节。</div><div class="notion-text notion-block-39e3936163f4815f9cacf0f54b2d3975">最后把固定点本征坐标和临界指数联系起来。这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 与 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 分别表示 thermal 与 magnetic relevant eigen-directions 上的耦合增量坐标；其他本征方向上的耦合增量记为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。它们在一次尺度因子为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 RG 下变为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d9a5f0ff8fa03551fe">现在看奇异自由能。设系统线度为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，体积为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。一次 RG 后，长度单位放大 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 倍，粗粒化后的系统线度为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，体积为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。奇异自由能总量描述同一个长距离临界涨落，因此满足</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814f8760f2a61e1fb040">把自由能密度定义为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b8b983e381a974f26a">约去 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，并代入本征坐标的缩放，就得到自由能密度的齐次标度律：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ef8e1ee0d6c4cd3ce5">相关长度的标度律也来自同一个 RG 变换。相关长度是一个长度量；一次尺度因子为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 RG 后，耦合坐标变为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而用新格距为单位表示的相关长度缩小 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 倍。因此等价地写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d89c16e409790786d8">下面从相关长度标度律推出相关长度指数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-39e3936163f4813c921bc192277eefaa">先看零外场，并只保留主导的 thermal relevant direction；irrelevant 坐标对主导奇异幂律只给出修正。因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4813ab5acdf208d52ed97">因为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是任意缩放因子，可以选它把右边第一个变量变成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4813f9c92c8ee5dcb2daa">右边的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是离临界点有限距离处的有限常数，所以主导奇异性由前面的幂次决定。相关长度临界指数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 定义为临界点附近相关长度的发散幂律：若 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是临界点两侧的非普适振幅，则 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就是这个幂律的指数。因此这里得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c4ac6ad297f040b3bd">同一选择代入自由能标度律。仍取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>、忽略 irrelevant 坐标的主导贡献，并取 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，则</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814ea891eef678a2ce61">因此奇异自由能密度的主导幂律为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481368d87f3ff46c02003">比热的奇异部分由自由能密度对温度变量的二阶导数给出。由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 求导只差一个非奇异常数因子，因此可以用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 求导来读出临界幂次：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481dcb3e8dcc1e829f9fa">另一方面，比热指数的定义是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c4b399de473b7bc36c">比较两个幂次，得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481cbaa3dc99506e2be5e">再用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到 hyperscaling 形式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ab8fe4e240e7ebd0c2">由 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 还得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481cb9f06f6c9a2f6fda5">若把 magnetic scaling dimension 写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4810b9eb4f1072a7ad91b">这些关系就把 Kadanoff 的固定点图像和通常的临界指数体系连接起来。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f481a18bd2dd7cb9994bc6" data-id="39e3936163f481a18bd2dd7cb9994bc6"><span><div id="39e3936163f481a18bd2dd7cb9994bc6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481a18bd2dd7cb9994bc6" title="Comparison with Landau Mean-Field Exponents"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Comparison with Landau Mean-Field Exponents</span></span></h4><div class="notion-text notion-block-39e3936163f481dc9d5ccca3b6cf8b90">Landau 理论给出的指数，是忽略临界涨落后得到的 mean-field exponents：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814f967af5fec8db87d3">在 RG 语言中，这组结果对应 Gaussian fixed point，也就是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的固定点。在线性化层面，Gaussian fixed point 给出</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481948815da0f56ff4cc1">如果暂时把 Gaussian fixed point 的本征值形式代入前面的 hyperscaling 型指数关系，会得到</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4813f917ce3348e84dcee">在上临界维数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 处，这个形式代入给出 Landau mean-field 指数</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481368f8ae270f1949daa">也就是 Landau mean-field 结果。这里要特别注意 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时 hyperscaling 给出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，也与 Landau 一致。</div><div class="notion-text notion-block-39e3936163f4812bbecff28d276c77ff">但是在 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 耦合是 relevant，Gaussian fixed point 不再控制临界点。RG 流被吸引到非平庸的 Wilson-Fisher fixed point，固定点处的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 被涨落修正，因此指数偏离 Landau 值。例如 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 Ising 普适类一阶结果为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481d39224e9cf1c3ba2d9">这就是二者的核心区别：Landau 理论等价于停在 Gaussian fixed point；真正低于上临界维数时，涨落把临界行为带到 Wilson-Fisher fixed point，临界指数由该非平庸固定点的本征值决定。</div><div class="notion-text notion-block-39e3936163f481d0bc21f3c618c3b015">对于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，Gaussian fixed point 控制临界行为，Landau 指数仍然成立；但 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 耦合虽然 irrelevant，却是 dangerously irrelevant variable。它不会改变固定点位置，却会进入磁化和自由能的奇异振幅，因此不能把前面的 hyperscaling 型公式机械地用于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；特别是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 一般不再适用。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-39e3936163f481ffabc8ca3fc6135495" data-id="39e3936163f481ffabc8ca3fc6135495"><span><div id="39e3936163f481ffabc8ca3fc6135495" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481ffabc8ca3fc6135495" title="2.4 Wilson momentum-shell RG"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2.4 Wilson momentum-shell RG</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f48101a9baef4fde538245" data-id="39e3936163f48101a9baef4fde538245"><span><div id="39e3936163f48101a9baef4fde538245" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f48101a9baef4fde538245" title="Setup and three-step RG procedure"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Setup and three-step RG procedure</span></span></h4><div class="notion-text notion-block-39e3936163f48154b490e08ccf9b401e">Wilson momentum-shell RG 的思想是：短距离涨落对应高动量模式，长距离涨落对应低动量模式。既然我们想知道长距离物理，就可以逐步积分掉高动量薄壳中的快模，再把系统重新缩放回原来的 cutoff。</div><div class="notion-text notion-block-39e3936163f481958ef1d9e450a62a3e">从带 ultraviolet cutoff 的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 理论出发：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4816e831ed8473a862415">with</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481fdb190c8d16d4a4397">这里的 cutoff 要按有效理论来理解。<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的自由度被定义为只包含 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 Fourier 模式；也就是说，在这个 effective field theory 内，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的模式不是独立积分变量。它们并不是“物理上不存在”，而是已经在从微观模型到粗粒化场的过程中被积分掉或平均掉，其影响被吸收到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 以及更高阶局域算符的系数里。</div><div class="notion-text notion-block-39e3936163f4810796a8fea49c2702e2">因此，若把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 看作当前有效理论的变量，则可以说 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 不属于该理论的变量空间，或者等价地在这个表示中取为零。但这只是有效场 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的定义，不是对微观自由度的断言。</div><div class="notion-text notion-block-39e3936163f481a6b0bad48cef45695f">写成 sharp cutoff <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是一种方便的理想化。真实的 coarse-graining 更常给出平滑的窗口函数，例如前面 block averaging 中的 form factor。只要关心的是长距离临界行为，sharp cutoff 与 smooth cutoff 的差别会改变非普适参数和高阶算符系数，但不改变普适临界指数。</div><div class="notion-text notion-block-39e3936163f481738f6adbe4c2fe525f">把场分解成慢模和快模：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48108aeffe2abd423b1c7">第一步是积分掉高动量薄壳中的快模：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481a4b33ee25d9075b840">这一步是 Kadanoff block spin 中“把块内部自由度求和”的连续版本。区别只是：这里被求和的不是某个实空间盒子里的自旋，而是动量空间薄壳 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中的短波涨落。</div><div class="notion-text notion-block-39e3936163f4816aafbde894c9b0991e">第二步是 rescale momentum，把 cutoff 恢复为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ebb09fe7c413b24c8a">这一步的意义是把一次 coarse-graining 后改变了的 cutoff 和长度单位恢复到原来的形式。积分掉 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的快模以后，剩下的慢模只满足 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。如果不 rescale，新的理论就定义在不同 cutoff 上，不能直接看作原来理论空间中的一个新点。</div><div class="notion-text notion-block-39e3936163f4812ebf83ed5537ea7b4b">通过 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，剩余区域 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 被重新映到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。这样 RG 的结果不表现为理论定义域、格距或 cutoff 的改变，而表现为同一 cutoff 下耦合常数和场归一化的改变。也就是说，重整化后的变化被吸收到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中，而不是留在系统结构本身的尺度单位里。</div><div class="notion-text notion-block-39e3936163f481f5a9f1ef8bc2cfd897">第三步是 rescale field，使梯度项保持标准归一化。若暂时忽略 anomalous dimension，在 tree-level 有</div><div class="notion-text notion-block-39e3936163f48105aebfddb7cdc052e7">令场的重标度先写成待定形式</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b2b14cce5337427114">由于 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4817ba7b7ece3ae613fc3">因此梯度项变为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4813a9119e9125c6fea9e">要求重标度后梯度项仍然保持标准归一化，即前面的系数仍为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，就必须有</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481de95d2daa3e0483a85">所以</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4811ea757e0e70f2421ee">经过这三步以后，新的有效泛函仍写成同样形式，但参数变了：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f481e39cdfe0f63a9c5791" data-id="39e3936163f481e39cdfe0f63a9c5791"><span><div id="39e3936163f481e39cdfe0f63a9c5791" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481e39cdfe0f63a9c5791" title="Tree-level calculation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Tree-level calculation</span></span></h4><div class="notion-text notion-block-39e3936163f481038b7cfe24772cc8b4">现在真实计算 tree-level 变换。tree-level 的意思是：先不计算快模积分产生的 loop correction，只保留 cutoff 恢复和场重标度带来的经典尺度因子。快模积分形式上仍写为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48142b662e96775ab2f40">在 tree-level 下，这一步只把 cutoff 从 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 降到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>；快模涨落对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的额外修正先不算。因此慢模部分写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4811383d5cd760b0aba59">接着恢复 cutoff。令</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4810ea308dd2235a9d32d">再取场重标度</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4813a9286ceb1a5be5988">这样梯度项保持标准归一化：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4818a829fedaa87da47be">质量项变为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481daa6efca2b2270f183">四次项变为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481deb376c846945f3807">外场项变为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48176bfcfc3f7ddd930dd">因此 tree-level Wilson RG 真实算出来是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48127baa4e7151f944e24">如果把 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，就可以写成 infinitesimal RG flow 的 tree-level 部分：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481218710db7efdb9402c">这立刻说明 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是 relevant perturbation，而 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 tree-level scaling dimension 是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 理论的上临界维数。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-39e3936163f481d2bfe3f6767dcb54ca" data-id="39e3936163f481d2bfe3f6767dcb54ca"><span><div id="39e3936163f481d2bfe3f6767dcb54ca" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39e3936163f481d2bfe3f6767dcb54ca" title="One-loop correction and Wilson-Fisher fixed point"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">One-loop correction and Wilson-Fisher fixed point</span></span></h4><div class="notion-text notion-block-39e3936163f481b994d4c43ed785e50a">现在把快模积分真正算到 one-loop。以下取 Ising 普适类对应的单分量实标量场，并使用前面的归一化</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481f99961fcb16daf1fcc">这里</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48117a9aec49c27fab9d4">把场分成慢模和快模</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48196ac4ef0625156523e">记快模 Gaussian 平均为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c0bf1cc021bef93f69">有效自由能由</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4817cb811c1781bd4d0ae">定义。对相互作用做 cumulant expansion：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481479c81dfb4ae33b77e">第一项 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 给出质量项修正。展开 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 中含两个慢模、两个快模的部分：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48111ab2bdcaf2eca9270">把它写成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到 tadpole correction</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c9a487d58aa1621679">第二个 cumulant 给出四次耦合的 one-loop 修正。保留四个外部慢模、内部两条快模收缩线的连通部分，有三个等价通道，因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481019aa9c57df1e6eddc">于是积分掉快模后、还未重标度前，慢模有效参数为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481ee8f93d0fde040cb30">现在取 infinitesimal shell，令 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示一次无穷小 RG 步长。定义</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48166993af3ffc541f834">薄壳积分在 leading order 为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481989b0cd047ee4a1f5b">引入无量纲耦合</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48138a931c97374d7d0c7">再做 momentum rescaling 和 field rescaling，使 cutoff 回到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，得到 one-loop RG 方程</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481b4b874ff1474b5fff1">在临界点附近 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，四次耦合的 beta function 变成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4815eac4ffea5638d879b">这就是前面写成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的结构；在当前归一化下 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><div class="notion-text notion-block-39e3936163f481fcbb80ff91fc8f9e10">固定点由 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 给出。除 Gaussian fixed point 外，还有非零固定点</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4814bbcfed7c34a396b42">这就是 Wilson-Fisher fixed point。它不是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的 Gaussian fixed point，而是由 tree-level 的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 放大效应和 one-loop 的 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 抑制效应平衡产生的非平庸固定点。</div><div class="notion-text notion-block-39e3936163f4814886cffa61237ea7eb">为了得到临界指数，在线性阶考察固定点附近的扰动。Jacobian 的两个本征值在 one-loop 阶为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481238686f63cb56622ae">第一个是 thermal relevant direction，第二个是 quartic coupling 方向在 Wilson-Fisher fixed point 附近的 irrelevant direction。于是</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f48191a93ef76eeb3ac557">在这个 one-loop 计算中，场的 anomalous dimension 仍为</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c88c17cea91dd1017c">因此</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f481c59510d7aadbebf5e1">再代入前面的指数关系，可以得到一阶 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 展开</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-39e3936163f4818a891fc7afb30e55cd">固定点坐标本身依赖 cutoff scheme 和耦合归一化；但本征值和由它们给出的临界指数是普适的。这就是 Wilson-Fisher fixed point 对 Landau 平均场结果的系统修正。</div><div class="notion-blank notion-block-39e3936163f480479d31e02166dbb40c"> </div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Chapter 4 循环神经网络]]></title>
            <link>https://blog.xiangsiqi.site/notes/DL4_cn</link>
            <guid>https://blog.xiangsiqi.site/notes/DL4_cn</guid>
            <pubDate>Mon, 27 Apr 2026 00:00:00 GMT</pubDate>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-8bb3936163f483afb1558111d862cb43"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-eb23936163f483e68bd3818d74bbb922" data-id="eb23936163f483e68bd3818d74bbb922"><span><div id="eb23936163f483e68bd3818d74bbb922" class="notion-header-anchor"></div><a class="notion-hash-link" href="#eb23936163f483e68bd3818d74bbb922" title="第一节 词向量"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">第一节 词向量</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-6303936163f4832193cc8195910f49ac" data-id="6303936163f4832193cc8195910f49ac"><span><div id="6303936163f4832193cc8195910f49ac" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6303936163f4832193cc8195910f49ac" title="词表示与 Embedding 矩阵"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">词表示与 Embedding 矩阵</span></span></h3><div class="notion-text notion-block-1293936163f483df89b90137a0f1cd9a">接下来需要把自然语言转换成模型可以处理的数字。第一步通常是 <b>tokenization（分词或切分 token）</b>。英文可以按空格和标点切分，中文则通常需要额外的分词工具，或者直接使用更细粒度的字、子词作为 token。例如：</div><div class="notion-text notion-block-2ce3936163f482b4850c81542fb30f5c">可以被切成：</div><div class="notion-text notion-block-86d3936163f48330b2e601133a0b4b76">得到所有 token 之后，就可以构建词表。词表本质上是一张从 token 到整数编号的映射表：</div><div class="notion-text notion-block-7963936163f48348bc7901bc2c46860d">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 用来补齐不同长度的序列，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示词表中没有出现过的未知词，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示句子开始，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示句子结束。这样一句话就可以被转换成编号序列：</div><div class="notion-text notion-block-5293936163f4836e91f7014b904f8ace">但是这里必须注意：这些数字只是词在词表中的<b>索引</b>，并不是真正有数值意义的特征。比如“我”被编号为 4，“喜欢”被编号为 5，并不表示“喜欢”比“我”大 1，也不表示它们在语义上距离更近；“学习”被编号为 7，也不表示它和“我”的语义距离是 3。编号的大小、差值、顺序都只是人为规定的词表位置，和词义没有直接关系。</div><div class="notion-text notion-block-e263936163f4827d8fa381e8c4a86dac">如果直接把这些编号当成普通数字输入神经网络，模型会误以为 0、1、2、3、4 之间存在连续的大小关系，就像年龄、温度、价格那样可以比较和相减。但词不是这样的对象。词与词之间的关系不是由编号差决定的，而是由语义、上下文和使用方式决定的。</div><div class="notion-text notion-block-df93936163f4834fba5001259b552e95">因此，我们需要把这些离散编号进一步转换成向量表示。最简单的方式是 one-hot 向量：如果词表大小为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，每个词就对应一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 维向量，其中只有该词对应的位置是 1，其余位置都是 0。这样做至少避免了“编号大小有意义”的误解。</div><div class="notion-text notion-block-6ab3936163f483cb9c59814c9b4a1506">从概率建模角度看，词表定义了模型每一步要预测的全部可能结果。如果词表大小为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，那么模型每个时间步最终通常输出一个 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 维向量，经过 softmax 后表示“下一个 token 是词表中每个词”的概率分布。</div><div class="notion-text notion-block-bd73936163f483a7808381ef444d5907">不过 one-hot 向量仍然非常稀疏，而且任意两个不同词之间都是正交的，无法表达语义相似性。所以现代神经网络中更常见的做法是通过 embedding 层，把离散编号映射成连续向量。这个连续向量才是模型真正用来计算的词表示。</div><div class="notion-text notion-block-7c23936163f482c193a80159ab35c9ee">例如，假设词表中有：</div><div class="notion-text notion-block-2193936163f48266aaaa0195da314d51">embedding 层并不是把 4、5、7 当成普通数字输入模型，而是把它们当成查表位置。假设 embedding 维度为 3，那么 embedding 矩阵可能是：</div><div class="notion-text notion-block-d663936163f482aca7fb8181d49c0e03">于是输入编号序列：</div><div class="notion-text notion-block-7663936163f48290ab1081e2900bf33b">真正进入模型的是：</div><div class="notion-text notion-block-f7f3936163f48297af7b81ef14f124fb">可以看到，编号只是用来找到对应行的地址；真正参与计算的是这一行向量。训练过程中，这些向量会被不断更新。如果“喜欢”和“学习”经常出现在相近上下文中，它们的向量就可能被学得更接近；如果两个词语义和用法差异很大，它们的向量距离就可能更远。</div><div class="notion-text notion-block-43e3936163f482b7813601c6a4275116">也可以构造一个 embedding 矩阵 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，写成</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-17d3936163f48347a6a0011f4d226ba5">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示 one-hot 向量，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示新的词特征向量。</div><div class="notion-text notion-block-9a63936163f482919bf901bf112c010d">在图像任务中，我们也可以说卷积网络在学习一种 embedding：它把原始像素逐层变成更抽象的特征向量。但图像的原始输入噪声很高，像素之间存在大量局部冗余，所以卷积层同时承担了去噪、局部特征提取和信息压缩的作用。</div><div class="notion-text notion-block-3293936163f482cebe798169e99169dd">语言任务有所不同。经过 tokenization 和词表映射之后，输入已经变成离散符号，例如“我”“喜欢”“学习”。这些符号不是原始像素那样的连续噪声信号，而是经过人类语言系统压缩后的单位。因此语言输入可以粗略看作一种低噪声、高抽象度的输入。正因为如此，我们不一定需要像图像那样先用卷积层逐级提取局部模式，而可以直接用一个 embedding 矩阵，把每个离散 token 映射成一个连续向量。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-5e33936163f48349bb088187c5537fc3" data-id="5e33936163f48349bb088187c5537fc3"><span><div id="5e33936163f48349bb088187c5537fc3" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5e33936163f48349bb088187c5537fc3" title="传统 embedding 训练方法"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">传统 embedding 训练方法</span></span></h3><div class="notion-text notion-block-6963936163f483d1a46b01fba77c66f9">Word2Vec、负采样、GloVe 和早期基于词向量的情感分类，都是传统 embedding 训练方法。它们在现代大模型中已经不再是 NLP 的主线，因为现代模型通常把 embedding 层和 Transformer 主体一起端到端训练，并且得到的是上下文化表示。但这些方法仍然有教学价值：它们解释了词向量为什么可以被学习、上下文为什么能塑造语义，以及大词表为什么会带来计算开销。</div><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-83e3936163f483f2b69381b647f873d4" data-id="83e3936163f483f2b69381b647f873d4"><span><div id="83e3936163f483f2b69381b647f873d4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#83e3936163f483f2b69381b647f873d4" title="Word2Vec"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Word2Vec</span></span></h4><div class="notion-text notion-block-3b03936163f48304821d81ca85402f5d">那么 embedding 矩阵如何学习？最自然的思路是设计一个预测任务，让模型在任务中自动学习词向量。典型结构如下：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-9453936163f4822fbcb481a5bb2a80bc"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ac17ee253-1fd0-478c-b731-e88829a4f76f%3Aimage.png?table=block&amp;id=94539361-63f4-822f-bcb4-81a5bb2a80bc&amp;t=94539361-63f4-822f-bcb4-81a5bb2a80bc" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-2fc3936163f48307be2f81b8df2fded0">这本质上是一个词预测任务。我们可以使用固定窗口，从任意长度的句子中采样训练样本。</div><div class="notion-text notion-block-8c93936163f482fbbbee8136f9493470">还可以设计不同的训练任务：</div><ul class="notion-list notion-list-disc notion-block-84d3936163f482b59618811737899240"><li>使用左右各 4 个词预测中心词，这就是<b>CBOW</b>。</li></ul><ul class="notion-list notion-list-disc notion-block-4de3936163f483b495d501ef123f67f8"><li>使用中心词预测附近词，这就是 <b>skip-gram</b>。</li></ul><h4 class="notion-h notion-h3 notion-h-indent-2 notion-block-2283936163f48215b9490184f5ad1013" data-id="2283936163f48215b9490184f5ad1013"><span><div id="2283936163f48215b9490184f5ad1013" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2283936163f48215b9490184f5ad1013" title="GloVe 词向量"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">GloVe 词向量</span></span></h4><div class="notion-text notion-block-97b3936163f48380919e01a723ae258a">GloVe（Global Vectors）的思想和 Word2Vec 不太一样。Word2Vec 是通过一个局部预测任务学习词向量，例如用中心词预测上下文；GloVe 则更直接：它先统计整个语料中词与词一起出现的次数，然后让词向量去拟合这些全局统计规律。</div><div class="notion-text notion-block-3883936163f4836088da81f30f30ec73">首先构造一个共现矩阵 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。矩阵中的元素 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示：当第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个词作为中心词时，第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个词在它附近窗口中出现了多少次。例如，如果我们在大量文本中发现“深度”附近经常出现“学习”，那么 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就会比较大；如果“深度”和“香蕉”几乎不一起出现，那么 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就会很小。</div><div class="notion-text notion-block-77d3936163f4820e833b0166a53d817c">构造方式可以写成：</div><div class="notion-text notion-block-9c43936163f48221adcd81fadbb0deda">这里的窗口大小是人为设定的，比如向左向右各看 5 个词。窗口越大，统计到的关系越偏向主题相关；窗口越小，统计到的关系越偏向语法和局部搭配。</div><div class="notion-text notion-block-fe23936163f483e5a33f018e642399b0">GloVe 的核心假设是：如果两个词经常在相似上下文中出现，它们应该有相似的词向量。更具体地说，词向量的内积应该能够反映词与词之间的共现强度。因此它希望：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-1693936163f483a9b3ad8150b74822f8">这里 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可以理解为词 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 作为上下文词时的向量，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可以理解为词 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 作为中心词时的向量，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是偏置项。左边是模型根据词向量计算出的分数，右边是语料统计出来的真实共现强度。</div><div class="notion-text notion-block-7123936163f483e6a166812e1e471000">为什么右边使用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，而不是直接使用 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>？因为词频差异通常非常大。比如某些高频词可能出现几十万次，而低频词只出现几次。如果直接拟合原始次数，高频词会支配整个训练过程。取对数之后，数量级被压缩，模型更关注“是否显著共现”而不是只追逐极端高频词。</div><div class="notion-text notion-block-1eb3936163f483d595b381cca982165f">于是 GloVe 的训练目标就是让所有词对的预测分数尽量接近真实的对数共现次数：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-bb63936163f483f5b75b014f64306b03">这个式子可以逐项理解：</div><ul class="notion-list notion-list-disc notion-block-8683936163f4836d91b3018c16d5a27b"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：中心词 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和上下文词 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的向量相似度。</li></ul><ul class="notion-list notion-list-disc notion-block-45c3936163f483c5b496810ae608602f"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：两个词各自的偏置，用来吸收词频本身带来的影响。</li></ul><ul class="notion-list notion-list-disc notion-block-dbc3936163f482f690c88188b00d4b78"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：真实语料中统计到的共现强度。</li></ul><ul class="notion-list notion-list-disc notion-block-55c3936163f4820fa4cd012225fa226f"><li>平方项：预测值和真实统计值之间的误差。</li></ul><ul class="notion-list notion-list-disc notion-block-3d03936163f483e09053010c56dec910"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：对所有词对都计算这个误差。</li></ul><ul class="notion-list notion-list-disc notion-block-4863936163f483aaaa7c01e42aadffcb"><li><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>：权重函数，用来决定某个词对的误差有多重要。</li></ul><div class="notion-text notion-block-99f3936163f4825a8f00812b009720e0">为什么还需要权重函数 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>？因为共现次数太小的词对往往噪声很大，比如两个词只偶然同框一次，不应该让它们强烈影响训练；但共现次数太大的词对也不能无限放大，否则高频词会压倒一切。因此 GloVe 使用一个截断的权重函数：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-1483936163f4838da9b5813a0df5c273">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 很小时，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 也小，说明这个词对的统计不太可靠，训练时权重较低；当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 增大到 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 之后，权重最多就是 1，不会继续无限增大。这样既降低了低频噪声的影响，也避免了高频词完全主导训练。</div><div class="notion-text notion-block-1dc3936163f482c9b8f8811aa2443889">如果 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，说明两个词从未在窗口中共现，此时 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 没有定义。实际训练中通常只对 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 的词对计算目标，或者让 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，使这些没有共现的词对不参与这一项损失。</div><div class="notion-text notion-block-e153936163f482b181ec01d7a498ace6">最后，GloVe 会得到两套向量：中心词向量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和上下文词向量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。实践中可以只使用其中一套，也可以把二者相加或平均作为最终词向量。直观上，GloVe 学到的是：哪些词经常共享相似上下文，哪些词在全局语料统计中具有相近位置。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-0f63936163f483548001812cfeb351fa" data-id="0f63936163f483548001812cfeb351fa"><span><div id="0f63936163f483548001812cfeb351fa" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0f63936163f483548001812cfeb351fa" title="词向量偏见"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">词向量偏见</span></span></h3><div class="notion-text notion-block-25f3936163f482d291c3811e428b7f12">词向量中的偏见来自训练语料本身的偏见。例如 “Man is to computer programmer as woman is to homemaker.” 这种类比就反映了语料中的性别偏见。因为词向量是从真实语料中学习出来的，如果语料中长期把某些职业、身份、行为和特定性别、种族或群体绑定在一起，词向量空间也会把这种统计关系保留下来。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-2993936163f48271b66a81ff54c1d607"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A34d04717-cfe4-4821-bdd4-9e93d8ad2857%3Aimage.png?table=block&amp;id=29939361-63f4-8271-b66a-81ff54c1d607&amp;t=29939361-63f4-8271-b66a-81ff54c1d607" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-7793936163f482589d3881eb06ad6e46">以性别偏见为例，词向量去偏通常分为三步。</div><div class="notion-text notion-block-2f83936163f4837285070132dcbfe0d6">第一步是找到偏见方向。比如可以用多组性别成对词来估计性别方向：</div><div class="notion-text notion-block-c743936163f482f69eca01c98f3f9578">这些差向量大致都指向“男性到女性”或“女性到男性”的方向。把这些方向综合起来，就可以得到词向量空间中的性别方向 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。粗略理解，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就是词向量空间里专门表示性别差异的那条轴。</div><div class="notion-text notion-block-3413936163f483dc899801cac955bf65">第二步是对性别中性词做 neutralize（中和）。有些词本身不应该带有性别方向，例如：</div><div class="notion-text notion-block-2af3936163f4825f87ce813c13946391">对于这类词，我们希望去掉它们在性别方向上的投影。设某个词向量为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，性别方向为单位向量 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，那么 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 在性别方向上的投影为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-98b3936163f483daabe901947f4adf8b">去掉这一部分后得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-17a3936163f483298078810d4aa5a2e9">这样做的意思是：保留这个词作为职业、身份或概念的主要语义，但尽量去掉它不应该携带的性别偏向。</div><div class="notion-text notion-block-3cc3936163f4835e848e81e232ef6667">第三步是对性别相关词做 equalize（等距化）。有些词本身确实和性别有关，例如：</div><div class="notion-text notion-block-7603936163f482a3b97b01af1202491e">这些词不能简单中和，因为它们的语义本来就包含性别信息。我们希望的是：它们在性别方向上保持对称，同时在其他语义方向上尽量一致。比如 “grandmother” 和 “grandfather” 都表示祖辈亲属，区别主要应该只在性别方向上，而不应该让其中一个离 “nurse” 更近、另一个离 “doctor” 更近。</div><div class="notion-text notion-block-eea3936163f4836d9c9a018bb59e8a6d">因此 equalize 的目标是让一组成对词围绕中性中心对称分布。直观地说，就是先找到它们共同的非性别语义中心，再让它们沿性别方向等距离地分开。这样可以保留必要的性别差异，同时减少不必要的社会偏见。</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-b3f3936163f4828ea26f8112e698e8b2" data-id="b3f3936163f4828ea26f8112e698e8b2"><span><div id="b3f3936163f4828ea26f8112e698e8b2" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b3f3936163f4828ea26f8112e698e8b2" title="第二节 循环神经网络"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">第二节 循环神经网络</span></span></h2><div class="notion-text notion-block-9713936163f4835bb03701d2d9007162">循环神经网络（Recurrent Neural Network, RNN）的思想可以追溯到 20 世纪 80 年代。它的核心目标是让神经网络处理序列数据：当前时刻的计算不仅依赖当前输入，也依赖过去时刻留下的隐藏状态。因此，RNN 很自然地被用于语言、语音和时间序列建模。</div><div class="notion-text notion-block-1753936163f4830e82f50171c284115e">但是普通 RNN 很快暴露出一个关键困难：当序列很长时，梯度需要沿时间反复传播，容易出现梯度消失或梯度爆炸，导致模型难以学习长程依赖。为了解决这个问题，Hochreiter 和 Schmidhuber 在 1997 年提出了 LSTM（Long Short-Term Memory）。LSTM 通过记忆单元和门控结构，让模型能够选择性地保留、遗忘和输出信息，从而更稳定地建模长距离依赖。</div><div class="notion-text notion-block-0ea3936163f48245a8d0016588d85683">GRU（Gated Recurrent Unit）则是在 2014 年左右提出的更简化的门控 RNN。它保留了门控控制信息流的思想，但结构比 LSTM 更简单，参数更少，训练也更轻量。可以粗略理解为：RNN 是最基本的循环结构，LSTM 是更强但更复杂的长程记忆结构，GRU 则是在能力和简洁性之间做折中的版本。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-4853936163f483a78806814ab9a52232" data-id="4853936163f483a78806814ab9a52232"><span><div id="4853936163f483a78806814ab9a52232" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4853936163f483a78806814ab9a52232" title="语言生成建模"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">语言生成建模</span></span></h3><div class="notion-text notion-block-94a3936163f4831daad38194cb5daaad">语言生成建模的目标是给一个语言分配概率。假设一个句子由 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 组成，我们真正想建模的是整个句子同时出现的概率：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d093936163f482f2af7d011c482642f1">根据概率论中的链式法则，联合概率可以拆成一连串条件概率的乘积：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-02e3936163f482f4aac301d071462acf">也就是说，并不是一次性判断整个句子，而是在每个位置上问：在已经看到前面所有词的情况下，下一个词是什么。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-6353936163f4824b86ac01076530776d" data-id="6353936163f4824b86ac01076530776d"><span><div id="6353936163f4824b86ac01076530776d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6353936163f4824b86ac01076530776d" title="循环神经网络"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">循环神经网络</span></span></h3><div class="notion-text notion-block-b713936163f4828994360114a404e378">循环神经网络（Recurrent Neural Network, RNN）直接从这个设计出发，它的结构如下图所示：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-77c3936163f483fc86a681d962a7397c"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ad41b6204-fa2b-4c2a-b99b-2f0919448e67%3Aimage.png?table=block&amp;id=77c39361-63f4-83fc-86a6-81d962a7397c&amp;t=77c39361-63f4-83fc-86a6-81d962a7397c" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-7bb3936163f48268aa7201a8fd60e316">如果把 RNN 用作语言模型，那么每个时间步的输出 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是一个词表上的分类softmax概率分布。</div><div class="notion-text notion-block-8b33936163f4834e8525818905201f12">在生成文本时，我们可以从这个概率分布中采样得到下一个词，再把这个词作为后续时间步的输入。这样不断重复，就可以生成一个完整序列，并且把所有概率乘起来得到：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-41b3936163f482ba9038016b80796656">RNN 的前向传播公式为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-5003936163f482828e4881e000f68852">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个时间步的隐藏状态，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是当前输入，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 是上一个时间步传来的记忆；<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 控制当前输入如何影响隐藏状态，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 控制历史状态如何传递，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 则把隐藏状态映射为当前输出。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-0bf3936163f48310ba9601e8ae0abee4" data-id="0bf3936163f48310ba9601e8ae0abee4"><span><div id="0bf3936163f48310ba9601e8ae0abee4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0bf3936163f48310ba9601e8ae0abee4" title="RNN 的反向传播"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">RNN 的反向传播</span></span></h3><div class="notion-text notion-block-e533936163f4838ba02e810ac45dcbe0">直观上，RNN 的反向传播如下图所示：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-5053936163f4833f911e0158d7077ce0"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A457ced20-121a-44d9-b231-86e73433a796%3Aimage.png?table=block&amp;id=50539361-63f4-833f-911e-0158d7077ce0&amp;t=50539361-63f4-833f-911e-0158d7077ce0" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-db33936163f482eab14d01ff93491d1d">其中红色箭头表示前向传播，蓝色箭头表示反向传播。我们定义每个时间步的损失函数和总损失为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-30b3936163f482cd9bd80173bfcd688d">这个反向传播过程通常称为 <b>BPTT（Backpropagation Through Time，沿时间反向传播）</b>。它的核心思想并不神秘：我们先把 RNN 按时间展开，把同一个循环单元复制成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 上的一串网络，然后对这个展开后的深层网络做普通反向传播。</div><div class="notion-text notion-block-63d3936163f4825f86ea016570e16f5f">区别在于，RNN 在每个时间步使用的是同一组参数，例如 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>、<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>、<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。因此这些参数的梯度不是只来自某一个位置，而是来自所有时间步的累加：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-28c3936163f48374bdd0010f1b79ac04">也就是说，第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个时间步产生的损失会沿着时间链条向前传播，影响更早的隐藏状态；而每一个时间步对共享参数造成的影响，最后会被加总起来更新同一组参数。</div><div class="notion-text notion-block-0d13936163f483b5bcc3818710846904">对于某个隐藏状态 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，它接收到的梯度通常有两部分：一部分来自当前时间步自己的输出损失，另一部分来自后续时间步传回来的梯度：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d593936163f482c0bd2e01629d535f73">如果任务是 many-to-one，例如情感分类，可能只有最后一个时间步有直接损失；但前面的时间步仍然会通过后续隐藏状态收到梯度。</div><div class="notion-text notion-block-3c63936163f483489f920129a3bd4dc2">这也解释了为什么普通 RNN 容易出现梯度消失和梯度爆炸。因为梯度要沿着时间反复乘以类似 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和激活函数导数的项：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-33a3936163f483fd833c81293c15f55b">当这些连乘项的模长长期小于 1，梯度就会迅速消失；长期大于 1，梯度就会爆炸。因此 BPTT 虽然形式上只是普通反向传播，但在长序列上会暴露出 RNN 的核心困难。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-b2a3936163f483788c0d0137927dfd18" data-id="b2a3936163f483788c0d0137927dfd18"><span><div id="b2a3936163f483788c0d0137927dfd18" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b2a3936163f483788c0d0137927dfd18" title="RNN 的各种用法"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">RNN 的各种用法</span></span></h3><div class="notion-text notion-block-ebe3936163f4829d89d181a8f4baef6c">不同任务对应不同的 RNN 输入输出结构：</div><ul class="notion-list notion-list-disc notion-block-fdc3936163f483f8875e81a91d85f3a9"><li>one to many：文本生成等任务</li></ul><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-e9e3936163f4838fa5e9015b34bb199e"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Afa65661d-ae03-4830-aae7-ea254c4821ce%3Aimage.png?table=block&amp;id=e9e39361-63f4-838f-a5e9-015b34bb199e&amp;t=e9e39361-63f4-838f-a5e9-015b34bb199e" alt="notion image" loading="lazy" decoding="async"/></div></figure><ul class="notion-list notion-list-disc notion-block-5673936163f483ddabad8162679ec036"><li>many to one：常用于情感分类、文本分类等任务</li></ul><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-c7a3936163f4833a8ed881051e9ec3b6"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Aac4f626c-41a7-46a6-8492-cf0515770c3f%3Aimage.png?table=block&amp;id=c7a39361-63f4-833a-8ed8-81051e9ec3b6&amp;t=c7a39361-63f4-833a-8ed8-81051e9ec3b6" alt="notion image" loading="lazy" decoding="async"/></div></figure><ul class="notion-list notion-list-disc notion-block-5b53936163f4835395bb01c11f1209f6"><li>many to many，且输入输出长度不同：常用于机器翻译等任务</li></ul><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-7383936163f482268c120191f032dfe7"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3Ab83a054e-9c4e-4221-8225-4ab62a453436%3Aimage.png?table=block&amp;id=73839361-63f4-8226-8c12-0191f032dfe7&amp;t=73839361-63f4-8226-8c12-0191f032dfe7" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-bfd3936163f4833ea54481ef32acc7d5" data-id="bfd3936163f4833ea54481ef32acc7d5"><span><div id="bfd3936163f4833ea54481ef32acc7d5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#bfd3936163f4833ea54481ef32acc7d5" title="双向 RNN"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">双向 RNN</span></span></h3><div class="notion-text notion-block-3d13936163f483ff96258145dbcf4831">双向 RNN 可以和前面提到的 CBOW 做一个类比：CBOW 用目标词左右两侧的上下文来预测中心词，而双向 RNN 也是同时利用当前位置左侧和右侧的信息。区别在于，CBOW 通常把窗口内的词向量做简单组合，而双向 RNN 会分别用正向递推和反向递推建模上下文。</div><div class="notion-text notion-block-2d93936163f48221b0050159635e315d">当一个任务不仅需要当前位置之前的信息，还需要当前位置之后的信息时，就可以使用双向 RNN。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-6553936163f48240b344810c4143c6ea"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A11c13416-839f-46a6-bd55-c5b0b0968202%3Aimage.png?table=block&amp;id=65539361-63f4-8240-b344-810c4143c6ea&amp;t=65539361-63f4-8240-b344-810c4143c6ea" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-f6d3936163f48333b56c01c20c48c769">在这类任务中，预测某个位置的输出时，可能同时需要上文和下文。双向 RNN 的做法是同时进行正向递推和反向递推：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-e423936163f482aa8c2a8125ffa424a3"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A28300a3c-0191-4b49-85b5-7c4f5c3468bb%3Aimage.png?table=block&amp;id=e4239361-63f4-82aa-8c2a-8125ffa424a3&amp;t=e4239361-63f4-82aa-8c2a-8125ffa424a3" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-53b3936163f483df91a301e198b52783" data-id="53b3936163f483df91a301e198b52783"><span><div id="53b3936163f483df91a301e198b52783" class="notion-header-anchor"></div><a class="notion-hash-link" href="#53b3936163f483df91a301e198b52783" title="Deep RNN"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Deep RNN</span></span></h3><div class="notion-text notion-block-1eb3936163f4824db96a014062cbdfba">前面讨论的是 RNN 在时间方向上展开得有多长，也就是序列长度的问题。但对于复杂任务，模型往往还需要在层数方向上变深。深层 RNN 可以理解为：每个时间步内部不只做一次循环单元计算，而是把多个 RNN 层堆叠起来，使低层提取局部序列特征，高层进一步组合更抽象的时序表示。其结构可以写成：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-2fb3936163f482dc96dc81ddceba22c8"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A822d63ab-68bb-4dbd-bf0b-48cd6cf3de54%3Aimage.png?table=block&amp;id=2fb39361-63f4-82dc-96dc-81ddceba22c8&amp;t=2fb39361-63f4-82dc-96dc-81ddceba22c8" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-1d33936163f483f5b614012ab17f7049">对于深层 RNN，计算单元可以记为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-c833936163f48301b5d48106340280f7" data-id="c833936163f48301b5d48106340280f7"><span><div id="c833936163f48301b5d48106340280f7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c833936163f48301b5d48106340280f7" title="梯度消失、GRU 与 LSTM"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">梯度消失、GRU 与 LSTM</span></span></h3><div class="notion-text notion-block-4bd3936163f483d0a2e48172f31ea1f3">对于深层 RNN 或长序列 RNN，梯度消失是非常严重的问题，因为它会让网络只能学习短程模式。一个典型例子是句子 “The cat, which already ate…, was full”。普通 RNN 很难长期记住主语是单数的 “cat”，可能会被中间复数成分干扰，从而输出 “The cat, which already ate…, were full”。为了解决这类长程依赖问题，研究者提出了 GRU 和 LSTM。</div><div class="notion-row"><a class="notion-bookmark notion-block-35e3936163f480caadd1dbd9ab7c5ff3" href="https://arxiv.org/abs/1412.3555" target="_blank" rel="noopener noreferrer"><div><div class="notion-bookmark-title">Empirical Evaluation of Gated Recurrent Neural Networks on...</div><div class="notion-bookmark-description">In this paper we compare different types of recurrent units in recurrent neural networks (RNNs). Especially, we focus on more sophisticated units that implement a gating mechanism, such as a long...</div><div class="notion-bookmark-link"><div class="notion-bookmark-link-icon"><img src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Ficons%2Fapple-touch-icon.png?table=block&amp;id=35e39361-63f4-80ca-add1-dbd9ab7c5ff3&amp;t=35e39361-63f4-80ca-add1-dbd9ab7c5ff3" alt="Empirical Evaluation of Gated Recurrent Neural Networks on..." loading="lazy" decoding="async"/></div><div class="notion-bookmark-link-text">https://arxiv.org/abs/1412.3555</div></div></div><div class="notion-bookmark-image"><img style="object-fit:cover" src="https://www.notion.so/image/https%3A%2F%2Farxiv.org%2Fstatic%2Fbrowse%2F0.3.4%2Fimages%2Farxiv-logo-fb.png?table=block&amp;id=35e39361-63f4-80ca-add1-dbd9ab7c5ff3&amp;t=35e39361-63f4-80ca-add1-dbd9ab7c5ff3" alt="Empirical Evaluation of Gated Recurrent Neural Networks on..." loading="lazy" decoding="async"/></div></a></div><ul class="notion-list notion-list-disc notion-block-3463936163f482a5952c0198374459dc"><li>GRU（Gated Recurrent Unit）</li></ul><div class="notion-text notion-block-ed63936163f4822db53381625909f807">在普通 RNN 中，我们只计算 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。但在 GRU 中，还要计算门控项：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-7b63936163f482b796fb01272929fa7c">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示 update。由于 sigmoid 函数的输出位于 $(0,1)$，所以有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-10f3936163f483a2bcf181b5832f3cf0">GRU 的作用是用额外的线性层和 sigmoid 门，决定当前时间步的新信息应该写入多少、旧记忆应该保留多少。它给 RNN 引入了一种动态稀疏性：不是每个时间步都必须完全改写记忆，而是由门控制信息流。</div><div class="notion-text notion-block-3b13936163f482bcbff88194d21eb7f5">完整 GRU 中还有另一个门：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-7c23936163f483349e838183511fcfb0">并且有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><ul class="notion-list notion-list-disc notion-block-ca33936163f4821dbc5181d96081ac11"><li>LSTM（Long Short-Term Memory）</li></ul><div class="notion-text notion-block-f723936163f482e88d3f81824ea9fca9">LSTM 可以看作更完整的门控 RNN。它显式区分细胞状态 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和隐藏状态 $a$，并使用三个门：更新门、遗忘门和输出门：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-3483936163f483be962f8112bee0b13b">单元内部的计算为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-0d83936163f4837d9c9c810672933fc4">其中更新门控制新信息写入多少，遗忘门控制旧记忆保留多少，输出门控制细胞状态中有多少信息暴露为隐藏状态。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-a0a3936163f482f5947a81a411809589"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:480px;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A24685a06-627b-430d-aa21-c44fdc675697%3Aimage.png?table=block&amp;id=a0a39361-63f4-82f5-947a-81a411809589&amp;t=a0a39361-63f4-82f5-947a-81a411809589" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-cc43936163f483f383888198417f115b">如果门的计算还依赖记忆单元 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，这种结构称为 peephole connection。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-65e3936163f48337b59281c5ef092662" data-id="65e3936163f48337b59281c5ef092662"><span><div id="65e3936163f48337b59281c5ef092662" class="notion-header-anchor"></div><a class="notion-hash-link" href="#65e3936163f48337b59281c5ef092662" title="Attention 模型"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Attention 模型</span></span></h3><div class="notion-text notion-block-fd83936163f48292993201490dfab90d">普通 RNN 最大的问题之一是：第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个时间步只能直接看到第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个时间步传给它的隐藏状态。这个隐藏状态本质上是对过去所有信息的压缩，而压缩过程可能丢失信息；如果序列很长，早期信息在反复传递中还可能被进一步削弱甚至损坏。</div><div class="notion-text notion-block-ee73936163f4831d96c981b185c84b0a">因此，与其只依赖一个已经被压缩过的状态量，不如让当前时间步直接向前连接到多个历史状态，并用一个权重决定应该重点看哪里。这就是 Attention 的基本动机：不要把所有历史信息强行塞进一个固定长度的向量，而是允许模型在需要时回头查看不同位置的信息。</div><div class="notion-text notion-block-56a3936163f48316aac0817c66e43af7">Attention 模型的核心思想是给解码器一个注意力权重，使它在生成每个词时，可以对输入序列中的不同位置分配不同关注程度。结构如下：</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-9fc3936163f483bb854e813987c90ce4"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:100%"><img style="object-fit:cover" src="https://www.notion.so/image/attachment%3A2fb15172-1157-49be-9bd9-b951db07c31e%3Aimage.png?table=block&amp;id=9fc39361-63f4-83bb-854e-813987c90ce4&amp;t=9fc39361-63f4-83bb-854e-813987c90ce4" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-d713936163f482aaba800110e7a092f4"><span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示在生成 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，模型应该对编码器特征 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 投入多少注意力。并且有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-ecd3936163f482e3ba7281f69d592710">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 可以由一个小型神经网络计算，它的输入通常是上一时刻解码器状态 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 和编码器特征 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>。</div><h2 class="notion-h notion-h1 notion-h-indent-0 notion-block-3303936163f48289acf68144e2effee7" data-id="3303936163f48289acf68144e2effee7"><span><div id="3303936163f48289acf68144e2effee7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3303936163f48289acf68144e2effee7" title="第三节 语言模型"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">第三节 语言模型</span></span></h2><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-a433936163f48236b6ab81e5ca96a961" data-id="a433936163f48236b6ab81e5ca96a961"><span><div id="a433936163f48236b6ab81e5ca96a961" class="notion-header-anchor"></div><a class="notion-hash-link" href="#a433936163f48236b6ab81e5ca96a961" title="大规模预训练"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">大规模预训练</span></span></h3><div class="notion-text notion-block-2273936163f4839cbdf00108c57a7075">前面讨论的 Word2Vec、GloVe 等方法，核心目标是先在大语料上训练一个词向量表，然后把这个词向量表迁移到下游任务中使用。这个思路在当时非常重要，因为很多下游任务的数据很少，单独训练一个模型很难学到足够好的语言表示。</div><div class="notion-text notion-block-97c3936163f4832389fc81b9497d2180">现代语言模型把这个思路推进了一步：不再只预训练一个静态的 embedding 矩阵，而是预训练整个神经网络。也就是说，embedding 层、注意力层、前馈网络层，以及最后的输出层，都会一起在大规模文本上训练。训练完成后，模型不只是知道每个词大致是什么意思，还学到了语法结构、上下文依赖、常识关联和复杂的生成能力。</div><div class="notion-text notion-block-d373936163f483518aa881c4ae620f41">这就是<b>大规模预训练（large-scale pretraining）</b>的基本思想：先在海量无标注文本上设计一个通用的自监督任务，让模型学习语言本身的统计规律；然后再把这个模型迁移到具体任务中，例如分类、问答、翻译、摘要、代码生成等。</div><div class="notion-text notion-block-1c73936163f48200945081935f800791">对于自回归语言模型，最常见的预训练任务就是预测下一个 token：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a0b3936163f4830f93d901d392dbcf30">训练时，模型看到前面的 token，输出词表上的概率分布，然后用真实的下一个 token 作为监督信号。这种监督信号不需要人工标注，因为文本天然具有顺序：每一个位置的下一个 token 都可以作为训练标签。因此，只要有足够大的语料，就可以构造出大量训练样本。</div><div class="notion-text notion-block-0dd3936163f4835ea06a01683270260f">和传统 embedding 方法相比，大规模预训练的关键区别在于：它学到的不是固定词向量，而是<b>上下文化表示</b>。同一个 token 在不同句子中会得到不同的内部表示。例如“苹果”在“我吃了一个苹果”和“苹果发布了新手机”中语义不同，现代语言模型可以根据上下文给出不同表示，而静态词向量通常只能给它一个固定向量。</div><div class="notion-text notion-block-fef3936163f4838f9c40812c4cbd48dc">因此，现代语言模型可以理解为：embedding 不再是最终目标，而只是模型的第一层输入表示；真正重要的是经过多层网络加工后的上下文表示，以及由这些表示产生的下一个 token 概率分布。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-b433936163f4839f98e60136cd8c7f6b" data-id="b433936163f4839f98e60136cd8c7f6b"><span><div id="b433936163f4839f98e60136cd8c7f6b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b433936163f4839f98e60136cd8c7f6b" title="预训练的损失函数"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">预训练的损失函数</span></span></h3><div class="notion-text notion-block-cf63936163f48261b26281d4861a86ee">对于自回归语言模型，预训练目标是让模型尽可能准确地预测下一个 token。假设一个训练序列为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-8c83936163f482c58ab9019b29763bc8">模型在第 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 个位置看到前面的 token：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9093936163f4825bbaed813f3b2284ed">然后输出一个词表上的概率分布：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a1e3936163f4823abfd48188000c82ba">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 表示模型参数。真实的下一个 token 是 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，因此我们希望模型分配给真实 token 的概率越大越好。</div><div class="notion-text notion-block-1c23936163f4829aa9fd014497acbc10">从最大似然估计的角度看，我们希望最大化整段文本的概率：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-f363936163f48333b4d20119d330e915">为了计算方便，通常取对数，把连乘变成求和：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-84a3936163f48231ac5a01967c4c8940">在深度学习里，我们通常把最大化问题改写成最小化损失函数，于是得到负对数似然损失：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-60e3936163f48226a25601e6d60491a2">如果对每个位置取平均，就是更常见的形式：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a803936163f48267bd2a81e5a6d4b2d8">这个损失也可以理解成交叉熵损失。因为真实标签通常是 one-hot 分布：真实 token 对应的位置概率为 1，其余 token 概率为 0；模型输出的是 softmax 概率分布。交叉熵会惩罚模型没有把足够高的概率分配给真实 token。</div><h3 class="notion-h notion-h2 notion-h-indent-1 notion-block-31f3936163f4821a9cdd01c05a485ca7" data-id="31f3936163f4821a9cdd01c05a485ca7"><span><div id="31f3936163f4821a9cdd01c05a485ca7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#31f3936163f4821a9cdd01c05a485ca7" title="模型生成"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">模型生成</span></span></h3><div class="notion-text notion-block-2a63936163f483f8813d81b35a28ee4a">训练完成后，语言模型就可以用于生成文本。生成过程本质上是重复同一个步骤：给定已经生成的上下文，模型输出下一个 token 的概率分布，然后我们从这个分布中选择一个 token，把它接到上下文后面，再继续预测下一个 token。</div><div class="notion-text notion-block-dbf3936163f4827b8adf81d25f0894be">形式上，如果当前已经有：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-d803936163f482488daa01c20bfad1f7">模型会输出：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-6583936163f4830c9a1b8166aefb68b2">也就是词表上所有 token 的概率。问题在于：拿到这个概率分布后，我们应该如何选择下一个 token？不同选择方式会产生不同的生成风格。</div><div class="notion-text notion-block-7b03936163f4824dab160146a3cf750e">最简单的方法是 <b>greedy decoding（贪心解码）</b>。它每一步都选择当前概率最高的 token：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-a8b3936163f482cc8826018033816c37">比如模型给出：</div><div class="notion-text notion-block-7d43936163f4838ab3a9015b8a32c75f">greedy 就会直接选择“学习”。这种方法简单、稳定、速度快，但问题是它只看当前一步最优，不保证整句话最优。很多时候，当前概率最高的词接下去可能导致后续句子变差。</div><div class="notion-text notion-block-80e3936163f483a3903e012ce3f9b864">更稳妥的方法是 <b>beam search（束搜索）</b>。它不是每一步只保留一个候选，而是保留概率最高的前 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 条候选序列，其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 称为 beam size。例如 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，模型每一步都保留目前总概率最高的 3 条路径，然后继续扩展它们。这样可以避免过早地把一些暂时不是第一名、但后续可能更好的序列丢掉。</div><div class="notion-text notion-block-38d3936163f48283b99c01d5b1ba092e">Beam search 的目标可以写成寻找概率最高的完整序列：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-df63936163f48293a6eb01ed3a39ebb6">实际计算时通常使用对数概率，避免许多小概率连乘造成数值下溢：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-9153936163f482818ff5814eee6ee725">Beam search 常用于机器翻译、摘要等希望输出稳定、确定性较强的任务。但它也有缺点：生成结果可能偏保守、缺少多样性，而且容易偏向短句，所以实践中常配合长度惩罚。</div><div class="notion-text notion-block-0193936163f482df956d012c2ca5b737">除了 greedy 和 beam search，生成时还经常调节 <b>temperature（温度）</b>。Temperature 不是改变模型参数，而是在 softmax 前调整 logits 的尺度。设模型输出的 logits 为 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span>，带 temperature 的 softmax 为：</div><span role="button" tabindex="0" class="notion-equation notion-equation-block"><span></span></span><div class="notion-text notion-block-8f73936163f48200b5b481d536a2a736">其中 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 就是 temperature。</div><div class="notion-text notion-block-9e13936163f48283a1d2012d36543cd7">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，保持原始分布不变。 当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，logits 会被放大，概率分布变得更尖锐，高概率 token 更容易被选中，生成结果更稳定、更保守。极端情况下，<span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 很小，效果会接近 greedy。</div><div class="notion-text notion-block-9ea3936163f4820182da01ac9ad0924b">当 <span role="button" tabindex="0" class="notion-equation notion-equation-inline"><span></span></span> 时，logits 会被压缩，概率分布变得更平坦，低概率 token 也更有机会被采样到，生成结果更随机、更有多样性，但也更容易出现不稳定或不合理的内容。</div><div class="notion-text notion-block-f203936163f483ccab72811600c49570">例如原始分布可能是：</div><div class="notion-text notion-block-53c3936163f4820483d88122a47c3432">降低 temperature 后，可能变成：</div><div class="notion-text notion-block-1da3936163f482cdbdac015c005873b2">升高 temperature 后，可能变成：</div><div class="notion-text notion-block-abf3936163f4823da3a80186fb0758da">因此，temperature 控制的是生成的随机性和多样性：低温更确定，高温更发散。实际使用中，严肃问答、代码生成通常倾向较低 temperature；创意写作、头脑风暴可以使用较高 temperature。</div></main></div>]]></content:encoded>
        </item>
    </channel>
</rss>