3 Preliminaries
3.1 Polynomials and measurements
For complex numbers \(\alpha ,\beta \in \mathbb {C}\), write \(\alpha \approx _\varepsilon \beta \) when \(|\alpha -\beta |\le \varepsilon \). In particular,
Fix a prime power \(q=p^t\). Write \(\omega =e^{2\pi i/p}\), and let \(\operatorname {tr}:\mathbb {F}_q\to \mathbb {F}_p\) denote the finite-field trace.
The finite-field trace is the map \(\operatorname {tr}:\mathbb {F}_q\to \mathbb {F}_p\) defined by
Let \(a \in \mathbb {F}_q\). Then
If \(a=0\), then \(\operatorname {tr}[{\boldsymbol {x}}\cdot a]=0\) for every \({\boldsymbol {x}}\), so the average is \(1\). If \(a\neq 0\), choose \(y \in \mathbb {F}_q\) such that \(\operatorname {tr}[a\cdot y]\neq 0\), and set
Translation invariance of the uniform distribution gives
Since \(\omega ^{\operatorname {tr}[y\cdot a]}\neq 1\), it follows that \(C=0\).
Let \(v \in \mathbb {F}_q^m\). Then
By linearity of the trace and independence of the coordinates,
Proposition 3.2 shows that the \(i\)-th factor is \(1\) when \(v_i=0\) and \(0\) otherwise, so the product is \(1\) exactly when \(v=0\).
Let \(\mathcal{P}(m,q,d)\) be the set of polynomials \(g \in \mathbb {F}_q[x_1,\dots ,x_m]\) whose degree in each variable is at most \(d\), viewed as functions \(\mathbb {F}_q^m \to \mathbb {F}_q\) via evaluation. (Over finite fields, distinct low-degree polynomials can induce the same function; we identify elements of \(\mathcal{P}(m,q,d)\) with their polynomial representatives, not their functional equivalence classes.)
Here, “individual degree \(d\)” means degree at most \(d\) in each coordinate. In particular,
Let \(g, h:\mathbb {F}_q^m \rightarrow \mathbb {F}_q\) be two distinct polynomials of total degree \(d\). Then
Apply the Schwartz–Zippel lemma to the nonzero polynomial \(g-h\), which has total degree at most \(d\).
If \(g,h \in \mathcal{P}(m,q,d)\) are distinct, then
The difference \(g-h\) is a nonzero polynomial of total degree at most \(md\). Apply Lemma 3.6.
Let \(g, g' : \mathrm{Polynomial}\, \mathrm{params}\) be two distinct full polynomial outcomes with individual degrees at most \(d\). Then
This packages the \(md/q\) loss term used in the mainFormal self-consistency cascade (inductive_step.tex, lines 119–133) and in comMain (commutativity-G.tex).
Transport along the equivalence between coded \(\mathbb {F}_q\) points and the underlying scalar function space, then apply Lemma 3.7.
Let \(f,h : \mathrm{AxisLinePolynomial}\, \mathrm{params}\) be two line-polynomial outcomes with distinct underlying polynomials. Then
The native univariate root-counting bound is \(d/q\); the statement deliberately pads this to the paper’s ambient \(md/q\) loss for Lemma 6.1.
Reindex coded line parameters to the scalar field, count the roots of the nonzero univariate polynomial \(f-h\), and use \(m\ge 1\) to pad \(d/q\) to \(md/q\).
Let \(\mathcal H\) be a Hilbert space and let \(\mathcal A\) be a set of outcomes. A submeasurement on \(\mathcal A\) is a family \(A=\{ A_a\} _{a \in \mathcal A}\) of Hermitian positive semidefinite operators on \(\mathcal H\) such that \(\sum _a A_a \le I\). It is a measurement when \(\sum _a A_a = I\), and it is projective when each \(A_a\) is an idempotent projection. (In the Lean formalization, outcome sets are assumed finite throughout.)
We write \(\mathrm{PolySub}(m,q,d)\) for the submeasurements indexed by \(\mathcal{P}(m,q,d)\), and \(\mathrm{PolyMeas}(m,q,d)\) for the corresponding measurements.
Let \(A=\{ A_a\} _{a \in \mathcal A}\) be a family of operators and let \(f\colon \mathcal A \to \mathcal B\). The post-processed family \(A_{[f(a)=b]}\) is defined by
Let \(A=\{ A_a\} _{a \in \mathcal A}\) be a family of operators and let \(f\colon \mathcal A \to \mathcal B\). Then
Consequently, if \(\{ A_a\} \) is a submeasurement, respectively a measurement, then \(\{ A_{[f(a)=b]}\} \) is again a submeasurement, respectively a measurement.
Each \(a \in \mathcal A\) contributes exactly once to the right-hand side, namely in the summand indexed by \(b=f(a)\).
In the notation \(A_{[f(a)=b]}\), the measurement outcome is the variable to which the function is applied. For example, if \(G=\{ G_g\} \in \mathrm{PolySub}(m,q,d)\) and \(u \in \mathbb {F}_q^m\), then
Here \(g\) is the measurement outcome, and \(u\) is the evaluation point.
Let \(A=\{ A_a^x\} _{a \in \mathcal A}\) be a submeasurement. Its completion is the measurement \(\widehat A\) with outcome set \(\widehat{\mathcal A}=\mathcal A \cup \{ \bot \} \) given by \(\widehat A_a^x=A_a^x\) for \(a \in \mathcal A\) and \(\widehat A_\bot ^x = I-\sum _a A_a^x\).
3.2 Consistency and state-dependent distance
Let \(\lvert \psi \rangle \in \mathcal H_{\mathrm A} \otimes \mathcal H_{\mathrm B}\). Let \(A=\{ A_a^x\} \) and \(B=\{ B_a^x\} \) be submeasurements with the same answer set, and let \(\mathcal D\) be a distribution on the question set. We write
when
A symmetric projective strategy \((\psi ,A,B,L)\) is \((\varepsilon ,\delta ,\gamma )\)-good if and only if the following hold on the relevant test distributions:
Each subtest accepts exactly when the two measurement outcomes agree. Because the relevant families are measurements, Lemma 3.18 rewrites those acceptance probabilities as consistency bounds.
If \(A=\{ A_a^x\} \) and \(B=\{ B_a^x\} \) are measurements, then
Expand the off-diagonal sum by writing \(\sum _{a \ne b} A_a^x \otimes B_b^x = \sum _b (I-A_b^x) \otimes B_b^x\).
Let \(\lvert \psi \rangle \in \mathcal H\), let \(A=\{ A_a^x\} \) and \(B=\{ B_a^x\} \) be families of operators on \(\mathcal H\), and let \(\mathcal D\) be a distribution on the question set. We write
when
In particular, if \(A^x_a \otimes I \simeq _{\varepsilon } I \otimes C^x_a\), the question is what hypothesis on \(A\) and \(B\) allows one to conclude that
Let \(A=\{ A_a^x\} \) and \(B=\{ B_a^x\} \) be measurements and let \(C=\{ C_a^x\} \) be a submeasurement. If \(A_a^x \otimes I \simeq _\delta I \otimes C_a^x\) and \(A_a^x \otimes I \approx _\varepsilon B_a^x \otimes I\), then
Rewrite the inconsistency with \(C\) as total mass minus diagonal overlap. The total mass is unchanged because \(A\) and \(B\) are measurements, and the diagonal overlap changes by at most \(\sqrt\varepsilon \) by Cauchy–Schwarz.
If \(A\) and \(B\) are measurements and \(A_a^x \otimes I \simeq _\delta I \otimes B_a^x\), then
If both measurements are projective, the converse also holds: if \(A_a^x \otimes I \approx _{2\delta } I \otimes B_a^x\) then \(A_a^x \otimes I \simeq _\delta I \otimes B_a^x\).
Expand the squared norm of \((A_a^x \otimes I - I \otimes B_a^x)\lvert \psi \rangle \) and use the diagonal-overlap formula from Lemma 3.18; projectivity turns the diagonal inequality into an equality, giving the converse.
The implication above does not extend to submeasurements. For example, if \(A_a^x=0\) for all \(a\), then \(A_a^x \otimes I \simeq _0 I \otimes B_a^x\), but
which is nonzero unless \((I \otimes B_a^x)\lvert \psi \rangle =0\) for all \(x\) and \(a\).
3.3 Consistency from state-dependent distance
Let \(\{ A^x_a\} \), \(\{ B^x_a\} \), and \(\{ C^x_{a,b}\} \) be matrices. Suppose that \(A^x_a \approx _\gamma B^x_a\) and that for all \(x\),
Then
Similarly, suppose that \((A^x_a)^\dagger \approx _\gamma (B^x_a)^\dagger \) and that for all \(x\),
Then
To prove 4, write the difference as
Cauchy–Schwarz bounds its magnitude by
and the two hypotheses bound these factors by \(1\) and \(\sqrt{\gamma }\) respectively. For 5, rewrite the difference as
and apply 4 to the adjoint families.
Let \(A = \{ A^x_a\} \), \(B = \{ B^x_a\} \), and \(C = \{ C^x_a\} \) be submeasurements such that \(A^x_a \approx _{\delta } B^x_a\). Then
The difference has magnitude
By Cauchy–Schwarz this is at most
The first factor is at most \(\sqrt{\delta }\) by hypothesis, and the second is at most \(1\) because \(C\) is a submeasurement.
Let \(\{ A^x_a\} \), \(\{ B^x_a\} \), and \(\{ C^x_{a,b}\} \) be matrices. Suppose that \(A^{x}_a \approx _\delta B^{x}_a\) and that for all \(x\) and \(a\),
Then
The error term is
Using \(\sum _b (C^{x}_{a,b})^\dagger (C^{x}_{a,b}) \le I\), this is bounded by
For vectors \(\lvert \psi _1 \rangle ,\dots ,\lvert \psi _k \rangle \),
First,
for all real numbers \(x_1,\dots ,x_k\). Apply 6 to \(x_i=\lVert \lvert \psi _i \rangle \rVert \) and combine it with the triangle inequality.
Suppose \((A_i)_a^x \approx _{\delta _i} (A_{i+1})_a^x\) for \(i=1,\dots ,k\). Then
Expand \((A_1-A_{k+1})\lvert \psi \rangle \) as a telescoping sum and apply Lemma 3.26.
The general operator-family theorem above is complemented in Lean by reusable binary and three-step helper lemmas for operator expectations, ‘qSDD‘, and ‘SDDRel‘. The three-step ‘SDDRel‘ helper is the specialization used in the Step 6 projectivization chain in Theorem 2.14; this remark is only a cross-reference and makes no separate proof-level completeness claim.
Suppose \(A_a^x \otimes I \simeq _\varepsilon I \otimes B_a^x\), \(C_a^x \otimes I \simeq _\delta I \otimes B_a^x\), and \(C_a^x \otimes I \simeq _\gamma I \otimes D_a^x\), where all four families are measurements. Then
Convert the two consistencies involving \(C\) to state-dependent distance, compose them by the triangle inequality, and transfer the result back to consistency by Lemma 3.20.
If \(A=\{ A_a^x\} \) and \(B=\{ B_a^x\} \) are measurements such that \(A_a^x \otimes I \simeq _\delta I \otimes B_a^x\), and \(f\) is a map on the answer set, then
Merging outcome classes can only decrease the off-diagonal mass.
Let \(A=\{ A_a^x\} \) be a submeasurement and \(B=\{ B_a^x\} \) a measurement such that \(A_a^x \otimes I \simeq _\gamma I \otimes B_a^x\). Then
where \(A^x=\sum _a A_a^x\). In particular,
For the first approximation in 7, bound
by the inconsistency between \(A\) and \(B\). The second approximation is identical, applied to
Then compose the two bounds with Lemma 3.27.
Suppose \(A=\{ A_a^x\} \) is a projective submeasurement satisfying
Then for every operator \(B\) with \(0 \le B \le I\),
We prove the first approximation in 9 in two steps. First,
This is obtained by applying Cauchy–Schwarz to the difference and using \(B^2 \le B \le I\) together with 8. For the second step, move the remaining left copy of \(A^{\boldsymbol {x}}_a\) across the bipartition in the same way; projectivity then collapses the sandwich and yields the first approximation in 9. The second approximation in 9 is the same argument with the rightmost factor \(A_a^{\boldsymbol {x}}\) omitted.
3.4 Strong self-consistency
Let \(A=\{ A_a^x\} \) be a submeasurement and \(P=\{ P_a^x\} \) a projective submeasurement such that \(A_a^x \otimes I \approx _\varepsilon P_a^x \otimes I\). Then
Since \(P\) is projective,
Apply Proposition 3.24 twice to replace \((P^{\boldsymbol {x}}_a)^2\) first by \(A^{\boldsymbol {x}}_aP^{\boldsymbol {x}}_a\) and then by \((A^{\boldsymbol {x}}_a)^2\), and conclude with \((A_a^{\boldsymbol {x}})^2 \le A_a^{\boldsymbol {x}}\).
Let \(\lvert \psi \rangle \) be permutation-invariant and let \(A=\{ A_a^x\} \) be a submeasurement. We say that \(A\) is \(\delta \)-strongly self-consistent when
If \(A\) is \(\delta \)-strongly self-consistent, then \(A_a^x \otimes I \simeq _\delta I \otimes A_a^x\). If \(A\) is a measurement, the converse also holds.
Rewrite the off-diagonal mass as the difference between the total mass of \(A\) and the diagonal overlap in the strong self-consistency inequality. When \(A\) is a full measurement, the total mass is the identity, so this inequality becomes an equality; hence the ordinary self-consistency defect and the strong self-consistency defect coincide, giving the converse.
If \(A\) is \(\delta \)-strongly self-consistent, then
If \(A\) is projective, the converse also holds.
Expanding the squared norm gives
The strong self-consistency bound controls the last line by \(2\delta \). If \(A\) is projective, then 11 is an equality, which gives the converse.
If \(A\) is \(\delta \)-strongly self-consistent and \(f\) is a function on its answer set, then
The same computation as in Lemma 3.36 gives
The post-processed diagonal term dominates \(\mathbb {E}_{{\boldsymbol {x}}} \sum _a \langle \psi \rvert A^{\boldsymbol {x}}_a \otimes A^{\boldsymbol {x}}_a \lvert \psi \rangle \), so strong self-consistency bounds 12 by \(2\delta \).
Let \(A\) be a \(\delta \)-strongly self-consistent submeasurement and let \(B\) be a submeasurement such that \(A_a^x \otimes I \approx _\varepsilon B_a^x \otimes I\). Then
Bound \(\langle \psi \rvert B \otimes I \lvert \psi \rangle \) from below by \(\mathbb {E}_{{\boldsymbol {x}}}\sum _a \langle \psi \rvert B^{\boldsymbol {x}}_a \otimes B^{\boldsymbol {x}}_a \lvert \psi \rangle \). Then apply Proposition 3.24 twice to replace this by \(\mathbb {E}_{{\boldsymbol {x}}}\sum _a \langle \psi \rvert A^{\boldsymbol {x}}_a \otimes A^{\boldsymbol {x}}_a \lvert \psi \rangle \), and finish with strong self-consistency of \(A\).
Let \(A\) be a \(\delta \)-strongly self-consistent submeasurement and let \(P\) be a projective submeasurement such that \(P_a^x \otimes I \approx _\varepsilon A_a^x \otimes I\). Then for every function \(f\),
First prove
Expanding the squared norm and using projectivity gives
Proposition 3.33 bounds the first term by \(\langle \psi \rvert A \otimes I \lvert \psi \rangle +2\sqrt{\varepsilon }\), and Proposition 3.24 together with strong self-consistency bounds the third term below by \(\langle \psi \rvert A \otimes I \lvert \psi \rangle -\delta -\sqrt{\varepsilon }\). Substituting into 14 yields 13. Now combine 13 with
Let \(A=\{ A_a\} \) be a \(\zeta \)-strongly self-consistent submeasurement. Then
Apply Cauchy–Schwarz to the two-sided overlap \(\sum _a \langle \psi \rvert A_a \otimes A_a \lvert \psi \rangle \) and compare it with the strong self-consistency lower bound.
Let \(A=\{ A_a\} \) be a measurement that is \(\zeta \)-strongly self-consistent, and let \(B=\{ B_a\} \) be a submeasurement such that \(A_a \otimes I \approx _\delta B_a \otimes I\). Writing \(B=\sum _a B_a\), one has
Since \(0 \le I-B \le I\), the left-hand side is at most \(1-\langle \psi \rvert B \otimes I \lvert \psi \rangle \). Compare \(\sum _a \langle \psi \rvert B_a^2 \otimes I \lvert \psi \rangle \) with \(\sum _a \langle \psi \rvert A_a B_a \otimes I \lvert \psi \rangle \) and then with \(\sum _a \langle \psi \rvert A_a^2 \otimes I \lvert \psi \rangle \) using the \(\approx _\delta \) hypothesis, and finish with Lemma 3.40.
Let \(A=\{ A_a\} \) be a measurement that is \(\zeta \)-strongly self-consistent, let \(B=\{ B_a\} \) be a submeasurement such that \(A_a \otimes I \approx _\delta B_a \otimes I\), and let \(C\) be the measurement obtained by adding the missing mass \(I-B\) to one distinguished answer \(a^*\). Then
By Lemma 3.26,
It remains to bound the residual term. Since \(0 \le I-B \le I\),
By Proposition 3.24,
Proposition 3.43 gives
Hence 15 implies
Substituting this into the first bound proves the claim.
This is the same statement as Lemma 3.40: for a \(\zeta \)-strongly self-consistent submeasurement \(A\),
Immediate from Lemma 3.40.