MemoryПамять - bitsбит - of bf16от bf16 LensesЛинз 0
Click to place a lens · drag to move · scroll to resize · click to remove · R clears Клик ставит линзу · тянуть - двигать · колесо меняет размер · клик убирает · R очищает Tap the field to place a lens · tap it again to remove Коснитесь поля, чтобы поставить линзу · второе касание убирает
The weights sit behind a glass. Lenses open over the blocks a request needs. Place a few.
Веса лежат за стеклом. Линзы открываются над теми блоками, которые нужны запросу. Поставьте несколько.
Precision to match the question
Точность - такая, какой вопрос
You add 2 + 2 without thinking; a theorem in differential equations takes a while. A model answers both at full power. Ask it how a word is spelled and it still reads every one of its weights, so that question costs as much as the hardest one it can solve. And in everyday use such questions are most of them.
2 + 2 вы считаете не задумываясь, а над теоремой из дифференциальных уравнений придётся посидеть. Модель считает оба на полной мощности. Спросили, как пишется слово, - она прочитала все свои веса до последнего, и вопрос стоил как самый сложный, какой она вообще способна решить. А в повседневной жизни таких вопросов большинство.
Quantization does spend less, and pays for it in the quality of the answers: every scheme of it decides the precision in pieces fixed in advance - the whole model, a layer, a channel, a token, an expert. The pieces are cut by the architecture, and all that is left to decide is how many bits each one gets. The question never enters into it.
Квантование умеет тратить меньше, и расплачивается за это качеством ответов: любая его схема раздаёт точность кусками, нарезанными заранее - вся модель, слой, канал, токен, эксперт. Куски задаёт архитектура, и решить остаётся только, сколько бит достанется каждому. Самого вопроса в этом решении нет.
FoQLens reads its weights unevenly: as a request travels through the model towards an answer, the precision rises on the blocks that request needs, and the rest is read coarsely. Which blocks those are the model works out itself - the field above shows such a layout on the real weights of Gemma 4 E2B.
FoQLens читает свои веса неравномерно: пока запрос идёт сквозь модель к ответу, точность поднимается на тех блоках, которые этому запросу нужны, а остальные читаются грубо. Какие это блоки, модель решает сама - поле наверху показывает такую раскладку на настоящих весах Gemma 4 E2B.
Mixture of Experts comes closest - the strongest architecture there is today, and what most large models are built on. Its weights are split into groups in advance, and for each token a router switches a few of them on while the rest are not computed at all. The split is decided during training and is the same for every request, the border is sharp - a group is on or off - and the precision is one and the same for every weight. Hence the waste: whatever many groups need is held inside each of them as its own copy, and the memory goes on duplicates. Here only the ladder of depths is fixed in advance: which blocks are read deeper is worked out on the request itself, and the depth changes by rungs.
Ближе всех к этому Mixture of Experts - сегодня самая сильная архитектура, на ней построена большая часть крупных моделей. Веса там заранее поделены на группы, и роутер для каждого токена включает несколько из них, а остальные не считаются вовсе. Деление задано при обучении и одно на все запросы, граница резкая - группа включена или нет, - а точность у всех весов одна и та же. Отсюда и расточительность: то общее, что нужно многим группам, лежит в каждой своей копией, и память уходит на дубли. У нас заранее задана только лестница глубин: какие блоки прочитать глубже, определяется на самом запросе, и глубина меняется ступенями.
The goals
Цели
The same weights run on any device and under any load, and open to full precision where the question needs it.
Одни и те же веса работают на любом устройстве и под любой нагрузкой, а полную точность открывают там, где она нужна вопросу.
On the device. One file of weights for a phone, glasses or a laptop: the precision follows the battery, the heat and the free memory, and what the question needs is read deepest.
На устройстве. Один файл весов на телефон, очки и ноутбук: точность подстраивается под заряд, нагрев и свободную память, а то, что нужно запросу, читается глубже всего.
The price of a request. A simple request needs small zones and little computation; a hard one lifts wider zones and takes longer. The bill follows the difficulty of the request.
Цена запроса. Простому запросу хватает маленьких зон, и вычислений на него уходит мало; трудный поднимает зоны шире и считается дольше. Счёт выходит по сложности запроса.
A model that does not fit. More weights than the machine has memory for: only what the request needs is loaded, the rest is read coarsely or not read at all.
Модель, которая не помещается. Весов больше, чем памяти на машине: в память поднимается только то, что нужно запросу, остальное читается грубо или не читается вовсе.
Agents and reasoning. The first pass runs mostly at base precision and gives a draft; the next ones read deeper where the draft shows it is needed.
Агенты и рассуждение. Первый проход идёт почти целиком на базовой точности и даёт черновик; следующие читают глубже там, где черновик показал, что это нужно.
Under load the quality gives way slowly. When the machine is busy or hot, the precision comes off the blocks the request does not need first, and the answer holds longer.
Под нагрузкой качество падает не сразу. Когда машина занята или горячая, точность снимают сначала с тех блоков, которые запросу не нужны, и ответ держится дольше.
Progress
Прогресс
| FactsФакты | ProofПруфы | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| A format of our own - the model is stored as one stack, a base with refinements over it, and any depth is read from the same file without repacking: 2, 4, 6 or 8 bits from one copy, 1.63 GiB less than bf16 on E2BСвой формат хранения - модель лежит одной стопкой, база и уточнения поверх неё, и любая глубина читается из того же файла без перепаковки: 2, 4, 6 или 8 бит из одной копии, на 1.63 ГиБ меньше bf16 на E2B | .refocustensors | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| The card reads a layout inside the pass - a CUDA kernel of ours takes every block of rows to its own depth straight from the bytes of the stored copy, inside the pass itself: a decoding step at a batch of 32 takes 28.8 ms against 422.6 unpacked, with bf16 at 20.4. The precision can therefore change from one pass to the next, with nothing reloadedВидеокарта читает раскладку прямо в проходе - наше ядро CUDA берёт каждый блок строк до своей глубины прямо из байтов хранимой копии: шаг декодирования при батче 32 занимает 28.8 мс против 422.6 с распаковкой, у bf16 - 20.4. Значит точность можно менять от прохода к проходу, ничего не перезагружая | kernels | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| A coarser model answers where the precise one refuses - on questions the full model does not know, the full model says the answer is not there in 12.1% of cases, and read coarsely it says so in 4.6%, while accepted answers rise from 13.4% to 25.2% - and the trend holds as the reading gets coarserГрубая модель отвечает там, где точная отказывается - на вопросах, которых полная модель не знает, она говорит «здесь ответа нет» в 12.1% случаев, а прочитанная грубо - в 4.6%, и принятых ответов становится больше: с 13.4% до 25.2%. Чем грубее чтение, тем сильнее тренд | E001 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| The base precision went from nothing to 76% - two-bit weights gave incoherent answers at first; over a calibrated base the same two bits keep 76.7% of what the model knows, and the whole ladder is read up from that baseБазовая точность поднята с нуля до 76% - на двух битах модель сперва несла бессвязицу; над калиброванной базой те же два бита держат 76.7% её знаний, и с этой базы читается вся лестница | E003 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| The address of a question is real and cheap to read - a paraphrase is recognised as the same question 92% of the time against 72% by its words, and the first 4-8 layers carry the same address as a full passАдрес вопроса существует и читается дёшево - пересказанный вопрос узнаётся как тот же в 92% случаев против 72% по одним словам, а первые 4-8 слоёв несут тот же адрес, что и полный проход | E004 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
A map exists, and it pays - over 103 hard questions, the ones only the eight-bit model answers, a map built for the question holds 82% of its answers on 52% of its memory. Over another hundred, half of them ordinary:
| E005, E006 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| The map belongs to its question - spoil the map at the same memory to the byte - another question’s, the same figure moved elsewhere, rungs dealt at random - and it gives 0.59-0.62 against 0.86Карта принадлежит вопросу - испортить её при той же памяти до байта - взять чужую, сдвинуть свою фигуру, раздать ступени наугад - и выходит 0.59-0.62 против 0.86 | E006 |
FoQZones, the mechanism
Описание механизма FoQZones
Every block of weights is read to its own depth. Most of the network stays at a base precision; the blocks a question puts to work are read deeper, up to the top of the ladder. A layout is exactly that: a depth per block. Two controls bound it - how much of the network rises above the base, and how far above it the highest blocks go - and what the reading costs follows from them.
Каждый блок весов читается на своей глубине. Большая часть сети остаётся на базовой точности, а блоки, которые задействует вопрос, читаются глубже, вплоть до верха лестницы. Раскладка - это и есть глубина на каждый блок. Ограничивают её две ручки: какая доля сети поднимается над базой и насколько высоко поднимаются самые верхние блоки; цена чтения выходит из них.
Where the depths come from is what the two variants differ in.
Чем два варианта различаются - откуда берутся эти глубины.
The layer decides for itself: it looks at what has just entered it and sets the depth of the blocks ahead. The network is read once, and the decision travels with the pass. This is the variant to build - it runs on the card already, the scores and the levels living on the device, a decoding step of 39.6 ms at a batch of 32 against 31.5 for a fixed mixed layout. One question is open, and it is the formula by which a layer picks whom to lift.
Слой решает сам: смотрит, что в него вошло, и назначает глубину блокам впереди. Сеть читается один раз, и решение едет вместе с проходом. Это приоритетный вариант - он уже работает на карте, очки и уровни лежат на устройстве, шаг декодирования 39.6 мс при батче 32 против 31.5 у неподвижной смешанной раскладки. Открыт один вопрос - сама формула, по которой слой выбирает, кого поднять.
at a layer the pass holds h_k, the state that entered it
the score s_c = ??? which formula - still being looked for; three tried below
1 activity s_c = ‖x_c‖² how strong a signal enters the unit
2 votes s_c = Σ a_b · κ(b → c) a unit votes for those it feeds; κ from the weights
3 draft s_c = ‖y_c(base)‖² · Δ_c² the layer read at base shows who carries the question
reduced ŝ_c = s_c · Δ_c² / w_c Δ_c the quantization step, w_c the weights of the block
level ℓ(c) = γ + |{ j ≥ 1 : ŝ_c ≥ λ · 16ʲ }| capped at κ = γ + ⌊g · (m − γ)⌋
by share the top f of the layer by ŝ rise; the rest stays at the base
на слое проход держит h_k - состояние, вошедшее в слой
очки s_c = ??? какая формула - ищется; ниже три, которые пробуем
1 активность s_c = ‖x_c‖² насколько сильный сигнал входит в единицу
2 голоса s_c = Σ a_b · κ(b → c) единица голосует за тех, кого кормит; κ - из весов
3 черновик s_c = ‖y_c(база)‖² · Δ_c² слой, прочитанный на базе, показывает, кто несёт вопрос
приведённая ŝ_c = s_c · Δ_c² / w_c Δ_c - шаг квантования, w_c - веса блока
уровень ℓ(c) = γ + |{ j ≥ 1 : ŝ_c ≥ λ · 16ʲ }| с потолком κ = γ + ⌊g · (m − γ)⌋
по доле верхние f слоя по ŝ поднимаются, остальное остаётся на базе
The whole map is worked out before the first layer, from the address the first layers read. The depths are then a shape: zones grown around the peaks of that address, each falling off to the base at its edge, by the rules below. There is a hint that this does not work: a map predicted from the address gave the same answers as the common map, the one used for every question (E005).
Вся карта считается до первого слоя, по адресу, который читают первые слои. Глубины тогда складываются в фигуру: зоны растут вокруг пиков этого адреса и спадают к базе на краю, по правилам ниже. Есть намёк, что так не выйдет: карта, предсказанная по адресу, дала те же ответы, что и общая карта - одна на все вопросы (E005).
ladder ℓ₀ < ℓ₁ < … < ℓ_m the levels this model can be read at; ℓ₀ = ZERO
base G = ℓ_γ any rung; floor in the scripts
distance d(a, b) between blocks; not chosen
centers c_i = ??? where a zone grows from, out of the address; not found
width D of the network in d
1 radius R = f · D f = 0: the center only; f = 1: the whole network
2 ceiling κ = γ + ⌊g · (m − γ)⌋ g = 0: the zone is the base; g = 1: the top rung
3 profile stops s_κ < … < s_(γ+1) where each ring ends, in radii; above 1, past the edge
4 lift ρ_i(b) = d(b, c_i) / R, L_i(b) = max(0, 1 − ρ_i(b) / s_last) linear for now
5 combine L(b) = min(1, Σ L_i) or max L_i
6 level ρ* = s_last · (1 − L(b)) the ring that ρ* falls into; outside every zone: G
7 memory bits = Σ w_b · bits(b) / Σ w_b
лестница ℓ₀ < ℓ₁ < … < ℓ_m уровни, на которых читается эта модель; ℓ₀ = ZERO
база G = ℓ_γ любая ступень; floor в скриптах
расстояние d(a, b) между блоками; не выбрано
центры c_i = ??? откуда растёт зона, из адреса; способ не найден
ширина D ширина сети в d
1 радиус R = f · D f = 0: только центр; f = 1: вся сеть
2 потолок κ = γ + ⌊g · (m − γ)⌋ g = 0: зона равна базе; g = 1: верхняя ступень
3 профиль ступени s_κ < … < s_(γ+1) где кончается каждое кольцо, в радиусах; выше 1 - за краем
4 подъём ρ_i(b) = d(b, c_i) / R, L_i(b) = max(0, 1 − ρ_i(b) / s_last) пока линейный
5 сведение L(b) = min(1, Σ L_i) или max L_i
6 уровень ρ* = s_last · (1 − L(b)) кольцо, в которое попал ρ*; вне всех зон: G
7 память бит = Σ w_b · бит(b) / Σ w_b
The rules name no level: the ladder comes from the model and the way it is stored. In full: docs/precision-regulator.md.
Уровни в правилах не названы: лестницу задаёт модель и то, как она хранится. Полностью: docs/precision-regulator.ru.md.
What is claimed
Что заявлено
| Id | ClaimУтверждение | |
|---|---|---|
| H0 | Topics separate in the model’s representationsТемы разделяются в представлениях модели | the measurement caught the formatзамер поймал формат вопроса |
| H1 | The mask of a query is separable by topic and concentrated - there is an address to read at allМаска запроса разделима по темам и сконцентрирована - адрес существует и читается | half measuredполовина измерена |
| H2 | Zones of related topics overlap more than zones of unrelated onesЗоны родственных тем перекрываются сильнее, чем зоны неродственных | waitsждёт |
| H3 | Precision laid out by the query beats uniform quantization at the same memoryТочность, разложенная по запросу, обходит ровное квантование при той же памяти | the ceiling is measuredпотолок измерен |
| H4 | A draft read mostly at base precision, then refined with sharper zones, ends better than the same model at native precision - where there are iterations: an agent or a model's reasoningЧерновик, прочитанный в основном на базовой точности, а затем уточнённый более резкими зонами, заканчивает лучше той же модели в исходной точности - там, где есть итерации: агент или рассуждение модели | the premise showed upпредпосылка проявилась |
| H5 | A network trained with zoning and read with FoQZones beats a Mixture of Experts trained the classical way on the same data, holding no more in memory at any momentСеть, обученная с зонированием и читаемая через FoQZones, обходит Mixture of Experts, обученную классически на тех же данных, не держа в памяти больше ни в какой момент | waitsждёт |
| H6 | How good asking for a shorter answer and coarsening the weights each are at holding knowledge in compressed form - and below which step coarsening slides into nonsenseНасколько хорошо держат знание в сжатом виде просьба ответить короче и огрубление весов - и ниже какой ступени огрубление сползает в бессмыслицу | waitsждёт |
The full list, with the experiments behind each: docs/hypotheses.md.
Полный список, с экспериментами за каждой гипотезой: docs/hypotheses.ru.md.
Roadmap
Дорожная карта
The ceiling is measured: a map built for the request answers on half the memory where an even reading has already lost those answers. A map the model builds for itself while it answers is not there yet.
Потолок измерен: карта под запрос отвечает за половину памяти там, где ровное чтение ответы уже теряет. Карту, которую модель строит сама в момент ответа, пока получить не вышло.
| StepШаг | What it settlesЧто решает | |
|---|---|---|
| 1 | The bench - what everything is measured on: one stored copy of the weights, a k-quant base with refinements, read at 2, 4, 6 or 8 bits, a decoding step as one CUDA graphСтенд - на чём всё считается: одна хранимая копия весов, база k-quant с уточнениями, чтение на 2, 4, 6 или 8 битах, шаг декодирования одним CUDA-графом | doneсделано2026-09-11 → 09-14 |
| 2 | A corpus of what the model knows - selected by its own answers in three regimes and frozen: 20,640 questions, 18,576 E2B-it knows and 2,064 it does notКорпус того, что модель знает - собран по её собственным ответам в трёх режимах и заморожен: 20 640 вопросов, из них 18 576 модель знает и 2 064 нет | doneсделано2026-09-13 → 09-15 |
| 3 | Uniform quantization on the corpus - the answers at D8, D6, D4 and D2 over a base calibrated with an imatrix: D8 keeps 97.5% of the full model's knowledge, D6 96.2%, D4 90.8%, D2 76.7% (51.4% over the bench's own base, E002), so the filter is tested from D2; where the full model refuses, a coarser one answers - the premise of H4 (E001)Ровное квантование на корпусе - сколько модель помнит, когда все веса читаются на одной глубине: на 8 битах 97.5%, на 6 - 96.2%, на 4 - 90.8%, на 2 - 76.7% (над собственной базой стенда 51.4%, E002). Там, где полная модель отказывается отвечать, грубая отвечает - на этом стоит H4 (E001) | doneсделано2026-09-15 → 09-19 |
| 4 | The filter - how the zones are built: the source of the scores, the distance between blocks, the reach and the profileФильтр - как строятся зоны: источник очков, метрика расстояния между блоками, охват и профиль | nowидётsinceс 2026-09-15 |
| 5 | The address of a query - which signal identifies the query rather than its wording: the hybrid of neuron activity and head energy reads a paraphrase as the same query in 0.917 of the cases against 0.717 for a bag of its tokens, and the first 4-8 layers at base precision give the same address as a full pass; the depth of the reading cannot be chosen from the queryАдрес запроса - какой сигнал узнаёт сам вопрос, а не слова, которыми он задан: активность нейронов вместе с энергией голов внимания узнаёт пересказ как тот же вопрос в 92% случаев против 72% по одним словам, и первых 4-8 слоёв для этого хватает. Насколько глубоко читать веса, по вопросу не угадать | doneсделано2026-09-19 |
| 6 | The precision map of a query - over 103 questions only the top rung answers, a map built for the query holds 82% of the right answers on 52% of the whole network's memory, where an even four-bit reading at 56% holds none and a six-bit one at 78% holds 4%; a common map that knows no question holds 0.544 at the same price. The maps are built by oraclesКарта точности запроса - на 103 вопросах, которые берёт только полная точность, карта под запрос держит 82% ответов за 52% памяти; ровное четырёхбитное чтение, стоя дороже, не держит ни одного, шестибитное - 4%. Общая карта, одна на все вопросы, при той же цене держит 54%: разница и есть цена знания о вопросе | doneсделано2026-09-20 |
| 7 | The map against an even reading - over 150 questions, 50 hard and 50 ordinary with 50 paraphrases: an ideal map takes 26 hard questions of 50 at 52% of the memory, where an even six-bit reading at 78% takes 7. The bounds of the scale are computed from the ladder; a search over 197 points found nothing better. Three spoilings at memory equal to the byte give 0.59-0.62 against 0.86, so the address is what decidesКарта против ровного чтения - на 150 вопросах, 50 трудных и 50 обычных плюс 50 пересказов: идеальная карта берёт 26 трудных из 50 за 52% памяти, ровное шестибитное чтение за 78% - семеро из тех же 50. Пороги точности считаются формулой, перебор 197 вариантов лучше не нашёл. Три порчи карты при той же памяти до байта дают 0.59-0.62 против 0.86 - решает адрес | doneсделано2026-09-21 |
| 8 | A map the model builds itself - the same gain without looking at the answer: from the address of a query, or decided as the pass runs. Neither does it yetКарта, которую модель строит сама - тот же выигрыш, но без подглядывания в ответ: по адресу запроса или решением по ходу прохода. Мост от адреса к карте уже пробовали - он ответил как общая карта; формула для решения по ходу прохода ещё не найдена | nowидётsinceс 2026-09-20 |
| 9 | The regulator answers to the machine - under load and heat the precision comes off what the request does not need firstРегулятор отвечает машине - под нагрузкой и при нагреве точность снимается сначала с того, что запросу не нужно | waitsждёт |
| 10 | Agents on the FoQLens model - a chain of a draft and refinements against the same model at native precision and uniform quantization at the same memoryАгенты на модели FoQLens - цепочка черновика и уточнений против той же модели в исходной точности и ровного квантования при той же памяти | waitsждёт |
The corpus
Корпус
A layout can only be measured where the model answers from memory: if the answer sits in the question itself, there is nothing to measure. So the questions are selected by measurement - a question stays if the full model answers it right in its own words. It never sees the options, in any regime, and three read the answer it gives:
- the SQuAD score - the match against the reference answer;
- the full model itself, marking the answer as an examiner would;
- Claude.
The corpus is frozen: 20,640 questions - 18,576 the model knows and 2,064 it does not.
Проверять карту есть смысл там, где модель отвечает по памяти: если ответ лежит в самом вопросе, мерить нечего. Поэтому вопросы отобраны замером: вопрос остаётся, если полная модель отвечает на него верно своими словами. Вариантов ответа она не видит ни в одном режиме, а сам ответ читают трое:
- оценка SQuAD - совпадение с эталонным ответом;
- сама полная модель - она проверяет ответ как экзаменатор;
- Claude.
Корпус заморожен: 20 640 вопросов - 18 576 модель знает и 2 064 не знает.
| RegimeРежим | Where the answer isГде ответ | What the zones should doЧто должны делать зоны |
|---|---|---|
| ReadingЧтение | in the passage - SQuAD v2в отрывке - SQuAD v2 | little: the knowledge came in with the promptмало: знание пришло вместе с промптом |
| KnowledgeЗнание | only in the weights - TriviaQA, NQ-open, ARC-Challenge and ARC-Easy without their optionsтолько в весах - TriviaQA, NQ-open, ARC-Challenge и ARC-Easy без вариантов | decide - the regime the idea is aboutрешают всё - ради этого режима идея и затевалась |
| Two stepsДва шага | split across two passages - HotpotQAразнесён по двум отрывкам - HotpotQA | matter, but less than in the secondважны, но меньше, чем во втором режиме |
If the zones help as much where the answer sits in the text of the question as where it is only in the weights, the mechanism is not doing what it claims. How a corpus is chosen: docs/corpus.md.
Если зоны помогают там, где ответ лежит прямо в тексте вопроса, ровно так же, как там, где он только в весах, значит механизм делает не то, что заявлено. Как выбирается корпус: docs/corpus.ru.md.