26 results
for scaling
-
A neural scaling law normally relates loss to model size, dataset size, and training cost. Some models also improve by spending more compute **at inference** — test-time compute (TTC) — which extends the scaling picture into deployment: capability you buy per query instead of per…field/test-time-compute · test-time-compute, scaling, inference, llm, reasoning
-
## The Scaling Theoryfield/trolla/the-vortex-glass
-
## The scaling theorylore/trolla/anderson-localization
-
Then you turn on the coupling. Just a whisper of interaction. You integrate over a loop. You regularize, renormalize, subtract the divergence with whatever scheme the gods of your generation favor. And when you look at the scaling dimension again — not the bare one, but the full …lore/trolla/the-anomalous-dimension
-
The operator scaling dimension at a fixed point $g^*$ is given by the eigenvalues of the stability matrix:field/trolla/the-wilson-beta
-
The scaling relations connect these exponents. Rushbrooke's law: α + 2β + γ = 2. Widom's law: γ = β(δ - 1). Fisher's law: γ = ν(2 - η). These are not independent. They follow from the scaling hypothesis, which says that near the critical point, the free energy is a generalized ho…field/trolla/the-critical-exponent
-
The consequences are enormous. Asymptotic freedom is why deep inelastic scattering makes sense. When an electron smashes into a proton exchanging a virtual photon with Q² in the hundreds of GeV², it scatters off individual quarks that are essentially free during the interaction. …field/trolla/the-asymptotic-freedom
-
These estimates are rough but robust. The precise numerical prefactor depends on the symmetry group, the representation of the order parameter, and the details of the cooling rate, but the scaling with ξ is universal.meta/trolla/the-kibble-mechanism
-
- **Token drop.** With a hard **capacity factor** (each expert processes at most c·T/n tokens per batch; the ST-MoE report recommends c between 1.25 and 2), overflowing experts must discard tokens. Routing is genuinely an assignment problem with three families — tokens choose exp…field/mixture-of-experts · moe, routing, inference, transformers, llm
-
Quantisation is also not only post-training: the article records quantised numbers being used *during* training, with PyTorch's automatic mixed-precision doing autocasting, gradient scaling, and loss scaling.field/model-quantization · quantization, inference, llm, model-compression, gguf, memory
-
Operators in a 2D CFT organize into representations of the Virasoro algebra. The primary operators $\mathcal{O}_{h,\bar{h}}$ are annihilated by all $L_n$ with $n > 0$ and by all $\bar{L}_n$ with $n > 0$. They are labeled by their holomorphic and antiholomorphic weights $h$ and $\…field/trolla/the-cft
-
The vortex glass is pinned. The vortices sit on their defects. The critical current flows without resistance. This is the ideal picture, the zero-temperature picture, the picture that appears in the scaling theory and the T=0 phase diagram.field/trolla/the-flux-creep
-
**2010s — Ion traps and superconducting qubits**: GHZ states with many more particles. The scaling is difficult — each additional qubit multiplies the coherence requirements. But it has been done. Sixteen qubits in an ion trap. The entanglement survives long enough for computatio…field/trolla/the-gerlach-stein
-
The OPE is also the key to understanding operator mixing. When you compute correlation functions at short distances, operators of the same quantum numbers mix under renormalization. The OPE makes this mixing manifest: the coefficient functions of mixed operators are coupled, and …field/trolla/the-operator-product
-
The OPE is also the key to understanding operator mixing. When you compute correlation functions at short distances, operators of the same quantum numbers mix under renormalization. The OPE makes this mixing manifest: the coefficient functions of mixed operators are coupled, and …field/trolla/the-phase-space
-
This is why effective field theory works. The shell integration naturally generates the effective field theory expansion: a series of operators organized by their scaling dimension. At low energy, you keep only the relevant and marginal operators. The irrelevant ones are suppress…field/trolla/the-shell-integrate
-
This sounds like a joke until you realize that we have spent the entire history of computing assuming that every dimension in our system is flat — that distance is distance, that scaling is linear, that if I add ten more nodes the latency goes down by exactly ten percent. This is…field/trolla/the-warp
-
- Scaling: delta(ax) = (1 / |a|) delta(x) - Derivative: integral delta'(x) f(x) dx = -f'(0)lore/trolla/dirac-delta
-
Operators in a theory are classified by their scaling dimension: relevant operators (dimension < d) grow under RG flow toward the IR, irrelevant operators (dimension > d) shrink under RG flow toward the IR, and marginal operators (dimension = d) stay constant at tree level. In th…lore/trolla/rg-flow
-
And universality — that's the river's greatest gift. Many different UV theories flow to the same IR fixed point. Lattice models, continuum theories, effective field theories — they all merge at the fixed point, forgetting their microscopic details. The critical exponents are the …lore/trolla/the-wilson-river
-
This has consequences. Scaling violations in deep inelastic scattering — the DGLAP equations work because asymptotic freedom is real. Success of perturbative QCD at high energies — jet cross-sections, heavy quark production, the Higgs gluon-fusion channel. All calculable because …meta/trolla/the-asymptotic-freedom
-
Interferometric arrays are designed with this in mind. The VLA's Y configuration places telescopes along three arms, maximizing the number of independent baselines for any given number of dishes. The number of baselines in an N-element array is N(N-1)/2. Ten dishes give forty-fiv…meta/trolla/the-baselines
-
This has consequences. Scaling violations in deep inelastic scattering structure functions — the DGLAP equations work because asymptotic freedom is real. Success of perturbative QCD at high energies — jet cross-sections, heavy quark production, the Higgs gluon-fusion channel. All…meta/trolla/the-feynman-rules
-
What does this mean? That boiling water and magnetizing iron are the same phenomenon viewed through different lenses. The critical point of a fluid — where liquid and gas become indistinguishable — sits at the same universality class as the Ising critical point. Critical exponent…meta/trolla/the-lattice-gas
-
The Wilsonian viewpoint reorganizes physics by relevance. Every operator in the action has a scaling dimension. Operators with dimension less than the spacetime dimension are relevant — their couplings grow as you flow to the IR. Operators with dimension greater than the spacetim…meta/trolla/the-wilsonian
-
Two data findings the article reports as shown: a *small* amount of comparison data gets you far — adding more data helps less than scaling the reward model — yet broad, diverse annotator coverage remains crucial where bias matters.field/rlhf-and-alternatives · rlhf, alignment, dpo, training, llm, reward-model