13 Quantum Entropy
This chapter develops trace distance, von Neumann entropy, and their basic properties for finite-dimensional quantum systems.
13.1 Trace norm
For a matrix \(A\in M_{D}(\mathbb {C})\), the singular values \(s_0(A),s_1(A),\ldots \) form a finitely supported family: only finitely many are nonzero. The Schatten one-norm of \(A\) is their sum, that is, the sum of the finitely many nonzero singular values:
This is the \(p=1\) case of the Schatten \(p\)-norm [ Wol12 , Chapter 8, Section 8.1 ] . The equivalent closed formula \(\lVert A\rVert _1=\sum _{i=0}^{D-1}s_i(A)\), summing over all \(D\) singular values including trailing zeros, is part of Theorem 13.1.4.
The trace norm of \(A\in M_{D}(\mathbb {C})\) is its Schatten one-norm:
Wolf records the equivalent formula \(\lVert A\rVert _1=\operatorname{tr}\lvert A\rvert \) [ Wol12 , Chapter 8, Section 8.1 ] ; this is Theorem 13.1.5.
The Schatten one-norm and trace norm are the sums over the finite support of the singular-value sequence. Equivalently, the trace norm is the sum over the indices below the rank of the represented linear map.
The singular-value sequence satisfies \(s_i(A)=0\) for all \(i\geq \operatorname{rank}(A)\), so summing over the support and summing over \(\{ 0,\ldots ,\operatorname{rank}(A)-1\} \) give the same value.
For every \(A\in M_{D}(\mathbb {C})\),
Also, \(\lVert A\rVert _{\operatorname{tr}}=0\) if and only if \(A=0\). Equivalently, \(\lVert A\rVert _{\operatorname{tr}}{\gt}0\) if and only if \(A\neq 0\).
The singular values satisfy \(s_i(A)\geq 0\), so \(\lVert A\rVert _{\operatorname{tr}}=0 \iff \forall i,\, s_i(A)=0 \iff A=0\).
For every \(A\in M_{D}(\mathbb {C})\), with \(\lvert A\rvert =\sqrt{A^\dagger A}\) the positive-semidefinite square root of \(A^\dagger A\),
This is the formula \(\lVert A\rVert _1=\operatorname{tr}[\lvert A\rvert ]\) of [ Wol12 , Chapter 8, Section 8.1 ] ; the trace of the positive-semidefinite matrix \(\lvert A\rvert \) is real.
The map \(T^\dagger T\), for \(T\) the linear map on \(\mathbb {C}^D\) represented by \(A\), is represented by \(A^\dagger A\), so the singular values of \(A\) are \(s_i(A)=\sqrt{\lambda _i}\) with \(\lambda _0,\ldots ,\lambda _{D-1}\) the eigenvalues of \(A^\dagger A\). Diagonalizing \(A^\dagger A=U\operatorname{diag}(\lambda _i)U^\dagger \) gives \(\lvert A\rvert =U\operatorname{diag}(\sqrt{\lambda _i})U^\dagger \), hence
where the last step uses that \(\lvert A\rvert \) is positive semidefinite, so \(\operatorname{tr}\lvert A\rvert \geq 0\) is real.
For every \(c\in \mathbb {C}\) and \(A\in M_{D}(\mathbb {C})\), \(\lVert cA\rVert _{\operatorname{tr}} =\lvert c\rvert \, \lVert A\rVert _{\operatorname{tr}}\). This is the homogeneity axiom for matrix norms [ Wol12 , Chapter 8, Section 8.1 ] .
From \((cA)^\dagger (cA)=\lvert c\rvert ^2\, A^\dagger A\) and uniqueness of the positive-semidefinite square root, \(\lvert cA\rvert =\lvert c\rvert \, \lvert A\rvert \). Linearity of the trace gives \(\operatorname{tr}\lvert cA\rvert =\lvert c\rvert \, \operatorname{tr}\lvert A\rvert \), and the claim follows from Theorem 13.1.5.
For all unitaries \(U,V\in M_{D}(\mathbb {C})\) and every \(A\in M_{D}(\mathbb {C})\),
The trace norm is thus unitarily invariant [ Wol12 , Chapter 8, Section 8.1 ] .
From \((UA)^\dagger (UA)=A^\dagger U^\dagger UA=A^\dagger A\) the absolute values agree: \(\lvert UA\rvert =\lvert A\rvert \). For the right factor, \((AV)^\dagger (AV)=V^\dagger (A^\dagger A)V\), and since \(V^\dagger \lvert A\rvert V\) is positive semidefinite with
uniqueness of the positive-semidefinite square root gives \(\lvert AV\rvert =V^\dagger \lvert A\rvert V\). Cyclicity of the trace then yields \(\operatorname{tr}\lvert AV\rvert =\operatorname{tr}(\lvert A\rvert \, VV^\dagger ) =\operatorname{tr}\lvert A\rvert \). For the two-sided case, \(\lVert UAV\rVert _{\operatorname{tr}} =\lVert UA\rVert _{\operatorname{tr}} =\lVert A\rVert _{\operatorname{tr}}\) by applying the right and left cases in turn.
For every \(A\in M_{D}(\mathbb {C})\), the trace norm is the sum of the square roots of the eigenvalues of the positive operator \(A^\dagger A\):
Since \(s_i(A)=\sqrt{\lambda _i(A^\dagger A)}\) by the definition of singular value, the finite singular-value expansion (3) gives
For every \(A\in M_{D}(\mathbb {C})\), with \(\lambda _0,\ldots ,\lambda _{D-1}\) the eigenvalues of the matrix \(A^\dagger A\), \(\lVert A\rVert _{\operatorname{tr}} =\sum _{i=0}^{D-1}\sqrt{\lambda _i}\).
Diagonalizing \(A^\dagger A=V\operatorname{diag}(\lambda _i)V^\dagger \) gives \(\lvert A\rvert =V\operatorname{diag}(\sqrt{\lambda _i})V^\dagger \), so \(\operatorname{tr}\lvert A\rvert =\sum _{i=0}^{D-1}\sqrt{\lambda _i}\), and the claim follows from Theorem 13.1.5.
For every \(A\in M_{D}(\mathbb {C})\),
The first inequality is the \(p=1\), \(p'=2\) case of [ Wol12 , Eq. (8.1) ] ; the second is the corresponding case of [ Wol12 , Eq. (8.7) ] . Both are used in the proof of trace-norm convergence toward asymptotic states.
Write \(s_0,\ldots ,s_{D-1}\geq 0\) for the singular values of \(A\). Then \(\lVert A\rVert _2^2=\sum _i s_i^2\) and \(\lVert A\rVert _1=\sum _i s_i\). The first inequality follows by expanding \((\sum _i s_i)^2\) and using nonnegativity. The second is the Cauchy–Schwarz bound \((\sum _i s_i)^2\leq D\sum _i s_i^2\).
For every \(A\in M_{D}(\mathbb {C})\),
and the maximum is attained: some unitary \(U\) satisfies \(\operatorname{tr}[A^\dagger U]=\lVert A\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Eq. (8.11) ] .
For the upper bound, let \(v_0,\ldots ,v_{D-1}\) be an orthonormal eigenbasis of \(A^\dagger A\) with eigenvalues \(\lambda _i\), and let \(U\) be unitary. Expanding the trace in this basis and applying the Cauchy–Schwarz inequality,
since \(\lVert Av_i\rVert ^2=\langle v_i,\, A^\dagger A\, v_i\rangle =\lambda _i\) and \(\lVert Uv_i\rVert =1\).
For attainment, the vectors \(w_i=Av_i/\sqrt{\lambda _i}\), taken over the indices with \(\lambda _i\neq 0\), satisfy
so they form an orthonormal family. Extend it to an orthonormal basis \((w_i)_i\) of \(\mathbb {C}^D\) and let \(U\) be the unitary with \(Uv_i=w_i\) for all \(i\). Then
where the indices with \(\lambda _i=0\) contribute \(0\) to both sides because \(\lVert Av_i\rVert ^2=\lambda _i=0\) forces \(Av_i=0\).
For all \(A,B\in M_{D}(\mathbb {C})\), \(\lVert A+B\rVert _{\operatorname{tr}} \leq \lVert A\rVert _{\operatorname{tr}}+\lVert B\rVert _{\operatorname{tr}}\). Together with homogeneity (Theorem 13.1.6) and definiteness (Theorem 13.1.4), this completes the norm axioms of [ Wol12 , Chapter 8, Section 8.1 ] for the trace norm.
Choose by Theorem 13.1.11 a unitary \(U\) with \(\operatorname{tr}[(A+B)^\dagger U]=\lVert A+B\rVert _{\operatorname{tr}}\). Splitting the trace and applying the upper-bound half of the same theorem to \(A\) and to \(B\) separately,
Let \(H\in M_{D}(\mathbb {C})\) be Hermitian with Jordan decomposition \(H=H^+-H^-\) into orthogonal positive parts, \(H^\pm \geq 0\) and \(H^+H^-=0\). Then \(\lVert H\rVert _{\operatorname{tr}}=\operatorname{tr}[H^+]+\operatorname{tr}[H^-]\).
The sum \(H^++H^-\) is positive semidefinite, and by orthogonality its square is
so \(H^++H^-=\sqrt{H^\dagger H}=\lvert H\rvert \). Hence \(\lVert H\rVert _{\operatorname{tr}} =\operatorname{tr}\lvert H\rvert =\operatorname{tr}[H^+]+\operatorname{tr}[H^-]\) by Theorem 13.1.5.
Let \(A\in M_{D}(\mathbb {C})\) be Hermitian with positive part \(A^+\). There is a matrix \(\Pi \) with \(0\leq \Pi \leq \mathbb {1}\), \(\Pi ^2=\Pi \), and \(\Pi A=A^+\), namely the orthogonal projection onto the support space of \(A^+\).
Diagonalize \(A=U\operatorname{diag}(\lambda _1,\ldots ,\lambda _D)\, U^\dagger \) and set
where \(\chi \) is the indicator function of \((0,\infty )\). Each claim is read off eigenvalue-wise: \(0\leq \chi \leq 1\) gives \(0\leq \Pi \leq \mathbb {1}\), \(\chi ^2=\chi \) gives \(\Pi ^2=\Pi \), and \(\chi (\lambda ) \lambda =\max (\lambda ,0)\) gives \(\Pi A=A^+\).
Let \(X,C\in M_{D}(\mathbb {C})\) with \(X\geq 0\). If \(C\geq 0\), then \(\operatorname{tr}[CX]\geq 0\); if \(C\leq \mathbb {1}\), then \(\operatorname{tr}[CX]\leq \operatorname{tr}[X]\).
The first bound is Lemma 9.5.1. For the second, \(\operatorname{tr}[X]-\operatorname{tr}[CX]=\operatorname{tr}[(\mathbb {1}-C)X]\geq 0\) by the same lemma, since \(\mathbb {1}-C\geq 0\).
If \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) is a positive linear map and \(H\in M_{D}(\mathbb {C})\) is Hermitian, then \(T(H)\) is Hermitian.
Write \(H=H^+-H^-\) with \(H^\pm \geq 0\). Then \(T(H)=T(H^+)-T(H^-)\) is a difference of positive semidefinite matrices, hence Hermitian.
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a positive linear map and let \(H\in M_{D}(\mathbb {C})\) be Hermitian. Then \(\operatorname{tr}[(TH)^+]\leq \operatorname{tr}[T(H^+)]\).
Let \(\Pi _+\) be the support projection of \((TH)^+\) from Lemma 13.1.14, so that \(\Pi _+ T(H) = (T(H))^+\) and \(0\leq \Pi _+\leq \mathbb {1}\). The estimate follows from the chain
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map and let \(H\in M_{D}(\mathbb {C})\) be Hermitian, with Jordan decompositions \(H=P_+-P_-\) and \(T(H)=Q_+-Q_-\). Then \(\operatorname{tr}[Q_+]\leq \operatorname{tr}[P_+]\).
Lemma 13.1.17 gives \(\operatorname{tr}[(TH)^+]\leq \operatorname{tr}[T(H^+)]\). The result follows because trace preservation gives \(\operatorname{tr}[T(H^+)]=\operatorname{tr}[H^+]\).
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map. Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\lVert T(H)\rVert _{\operatorname{tr}}\leq \lVert H\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Theorem 8.16 ] .
Write \(H=P_+-P_-\) and \(T(H)=Q_+-Q_-\) for the Jordan decompositions. Applying Lemma 13.1.18 to \(H\) and to \(-H\) gives \(\operatorname{tr}[Q_+]\leq \operatorname{tr}[P_+]\) and \(\operatorname{tr}[Q_-]\leq \operatorname{tr}[P_-]\), so by the Jordan trace-norm formula (Lemma 13.1.13),
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map. Then for all density matrices \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\), \(\lVert T(\rho _1)-T(\rho _2)\rVert _{\operatorname{tr}} \leq \lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Eq. (8.80) ] .
By linearity \(T(\rho _1)-T(\rho _2)=T(\rho _1-\rho _2)\), and \(\rho _1-\rho _2\) is Hermitian as a difference of positive semidefinite matrices, so Theorem 13.1.19 applies.
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a complex-linear map. Then
The left supremum is over distinct density matrices in \(M_{D}(\mathbb {C})\), and the right supremum is over orthogonal unit vectors in \(\mathbb {C}^D\). No positivity or trace-preservation assumption is imposed on \(T\). This is [ Wol12 , Chapter 8, Lemma 8.3, Eq. (8.81) ] .
An orthogonal pair of pure states gives a pair of distinct density matrices at trace-norm distance \(2\), which proves the lower bound. For the reverse inequality, set \(H=\rho _1-\rho _2\) and write its Jordan decomposition as \(H=H^+-H^-\). Since \(H\) is nonzero and traceless, \(\operatorname{tr}[H^+]=\operatorname{tr}[H^-]=t{\gt}0\). Thus \(P=H^+/t\) and \(Q=H^-/t\) are density matrices with orthogonal supports, and homogeneity reduces the quotient for \(H\) to \(\frac12\lVert T(P-Q)\rVert _{\operatorname{tr}}\).
Take spectral decompositions \(P=\sum _i\lambda _i|\psi _i\rangle \! \langle \psi _i|\) and \(Q=\sum _j\mu _j|\phi _j\rangle \! \langle \phi _j|\). Orthogonality of the supports gives \(\psi _i\perp \phi _j\) whenever \(\lambda _i\mu _j\neq 0\), while \(\lambda _i,\mu _j\geq 0\) and \(\sum _i\lambda _i=\sum _j\mu _j=1\). The product-weight expansion
is therefore a convex combination of differences of orthogonal pure states. The triangle inequality and homogeneity bound its image under \(T\) by the right-hand supremum. Taking the supremum over \(\rho _1\neq \rho _2\) proves the equality.
For \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\), let \(T_\phi \) be its peripheral spectral projection and define
This is Wolf’s notation in [ Wol12 , Chapter 8, Proposition “Convergence towards asymptotic states”, Eq. (8.112) ] .
Let \(T_\varphi :=T\circ T_\phi =T_\phi \circ T\) be Wolf’s asymptotic dynamics. For every matrix \(\rho \), every \(n\geq 0\), and every \(n{\gt}0\) in the second equality,
These are the numerator identities used in [ Wol12 , Chapter 8, Eq. (8.114) ] . They follow from \(T_\phi T=TT_\phi =T_\varphi =T_\varphi T_\phi \); the positive-iterate condition records that the zeroth power of \(T_\varphi \) is the identity.
The commutation \(T_\phi T=TT_\phi \) iterates to \(T_\phi T^n=T^nT_\phi \), which proves the first equality by linearity. Idempotence of \(T_\phi \) gives \(T_\varphi (\rho -T_\phi (\rho ))=0\); every positive power of \(T_\varphi \) therefore vanishes on the same remainder, proving the second equality.
Let \(S:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be complex-linear and let \(\rho _1\neq \rho _2\) be density matrices. Then
This is the pointwise inequality from Wolf Lemma 8.3 used in Equation (8.115).
The proof is the upper-bound half of Lemma 13.1.21: normalize the positive and negative parts of \(\rho _1-\rho _2\) to orthogonally supported density matrices and expand both in eigenprojectors.
For \(A\in M_{D}(\mathbb {C})\) and a complex-linear map \(S:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\),
If \(\psi \) and \(\phi \) are orthogonal unit vectors, then \(\lVert |\psi \rangle \! \langle \psi |-|\phi \rangle \! \langle \phi |\rVert _2=\sqrt2\). Consequently the orthogonal-pure-state supremum for \(S\) is at most \(\sqrt D\, \lVert S\rVert _{2\to 2}\sqrt2\). These are precisely the estimates in Wolf Equations (8.115)–(8.116).
The first inequality is finite Cauchy–Schwarz applied to the singular values. The second is the defining application bound for the operator norm after Frobenius vectorization. Orthogonality makes the two rank-one projectors multiply to zero, so the squared Hilbert–Schmidt norm of their difference is the sum of their traces, namely \(2\).
Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be positive and trace preserving, let \(\rho \) be a density operator, and let \(n{\gt}0\). Then
Here the displayed transfer-matrix norm equals \(\lVert T^n-T_\varphi ^n\rVert _{2\to 2}\), the Hilbert–Schmidt operator norm of the corresponding superoperator. The explicit \(n{\gt}0\) records the paper’s positive-natural convention: at \(n=0\), both superoperator powers are the identity, so the displayed right-hand side vanishes. This is the upper assertion of [ Wol12 , Chapter 8, Proposition “Convergence towards asymptotic states”, Eqs. (8.112), (8.114)–(8.116) ] .
Apply the numerator identity to \(S=T^n-T_\varphi ^n\). The peripheral projection is again positive and trace preserving, so Lemma 8.3 applies to the distinct density matrices \(\rho \) and \(T_\phi (\rho )\). Then use \(\lVert \cdot \rVert _{\operatorname{tr}}\leq \sqrt D\lVert \cdot \rVert _2\) and the \(\sqrt2\) Hilbert–Schmidt distance of orthogonal pure states. The constants simplify as \(\frac12\sqrt D\sqrt2=\sqrt{D/2}\); if \(\rho =T_\phi (\rho )\), both sides vanish directly. Finally, the positive-power identity identifies the transfer-matrix difference with the transfer matrix of \(T^n-T_\varphi ^n\), whose largest-singular-value norm equals the Hilbert–Schmidt operator norm \(\lVert T^n-T_\varphi ^n\rVert _{2\to 2}\).
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a positive linear map with \(\operatorname{tr}[T(\rho )] = c\, \operatorname{tr}[\rho ]\) for some nonnegative real \(c\) and all \(\rho \in M_{D}(\mathbb {C})\). Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\operatorname{tr}[(TH)^+]\leq c\cdot \operatorname{tr}[H^+]\). Generalizes Lemma 13.1.18 from the trace-preserving case \(c=1\).
Lemma 13.1.17 gives \(\operatorname{tr}[(TH)^+]\leq \operatorname{tr}[T(H^+)]\). The trace-scaling hypothesis then yields \(\operatorname{tr}[T(H^+)] = c\, \operatorname{tr}[H^+]\), giving the claimed bound.
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a positive linear map with \(\operatorname{tr}[T(\rho )] = c\, \operatorname{tr}[\rho ]\) for some nonnegative real \(c\) and all \(\rho \in M_{D}(\mathbb {C})\). Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\lVert T(H)\rVert _{\operatorname{tr}}\leq c\cdot \lVert H\rVert _{\operatorname{tr}}\). Generalizes Theorem 13.1.19 from the trace-preserving case \(c=1\).
Let \(T,T':M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be trace-preserving and Hermiticity-preserving linear maps with \(T'(X)=\operatorname{tr}[X]\, Y\) for some fixed \(Y\in M_{D'}(\mathbb {C})\). If \(T-\varepsilon T'\) is positive for some \(\varepsilon \geq 0\), then for all density operators \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\),
See [ Wol12 , Chapter 8, Theorem 8.17 ] .
The map \(S:=T-\varepsilon T'\) is positive by hypothesis. Since \(T'(\rho _1-\rho _2)=\operatorname{tr}[\rho _1-\rho _2]\, Y=0\) (the densities have unit trace), \(T(\rho _1)-T(\rho _2)=S(\rho _1-\rho _2)\). The trace of \(S(\rho )\) is \((\operatorname{tr}[\rho ])-\varepsilon \operatorname{tr}[\rho ]=(1-\varepsilon ) \operatorname{tr}[\rho ]\), so \(S\) satisfies the scaled-trace hypothesis of Theorem 13.1.28 with \(c=1-\varepsilon \). Positivity of \(S\) on \(\rho _1\) forces \(0\leq 1-\varepsilon \) (by PSD trace nonnegativity), so the scaled-trace theorem applies and yields the claimed bound.
For any matrix \(Y\in M_{D'}(\mathbb {C})\), define \(T'_Y:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) by
For arbitrary \(Y\in M_{D'}(\mathbb {C})\), the rectangular (output-factor-first) Choi matrix of \(T'_Y\) is
See [ Wol12 , Chapter 8, Eq. (8.86) ] .
The normalized omega slice is \(D^{-1}E_{ij}\), where \(E_{ij}\) is the matrix unit. Since \(\operatorname{tr}(E_{ij})=\delta _{ij}\), the trace of the slice is \(D^{-1}\delta _{ij}\), and \(T'_Y\) multiplies this scalar by \(Y\). On the other hand, \((Y\otimes \mathbb {1})_{(a,i),(b,j)} = Y_{ab}\, \delta _{ij}\), so both sides equal \(D^{-1}Y_{ab}\, \delta _{ij}\).
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving, Hermiticity-preserving linear map and \(Y\in M_{D'}(\mathbb {C})\) Hermitian with \(\operatorname{tr}[Y]=1\). For \(\varepsilon \geq 0\), if the rectangular Choi matrix \(\tau (T)\) satisfies
then for all density operators \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\),
See [ Wol12 , Chapter 8, Eq. (8.86) ] .
Since \(\operatorname{tr}[Y]=1\), \(T'_Y\) is trace-preserving: \(\operatorname{tr}[T'_Y(X)]=\operatorname{tr}[X]\operatorname{tr}[Y]=\operatorname{tr}[X]\). Since \(Y\) is Hermitian and \(\operatorname{tr}[X]\) is real for Hermitian \(X\), \(T'_Y\) is Hermiticity-preserving. By Lemma 13.1.31, \(\tau (T'_Y)=(1/D)(Y\otimes \mathbb {1})\). Linearity of the Choi assignment gives \(\tau (T-\varepsilon T'_Y)=\tau (T)-\varepsilon \tau (T'_Y) \ge 0\), so \(T-\varepsilon T'_Y\) is completely positive (the rectangular Choi CP equivalence, Theorem 2.6.6) and hence positive. Theorem 13.1.29 then yields the asserted bound.
13.2 Projective pinching
Let \(\rho \in M_{d}(\mathbb {C})\) be a density matrix and let \(P_1,\ldots ,P_k\in M_{d}(\mathbb {C})\) be orthogonal projections satisfying \(\sum _{i=1}^k P_i=\mathbb {1}\). Define
Then
This is [ Wol12 , Chapter 8, Eq. (8.56) ] .
For \(A_{ij}=P_i-P_j\), positivity of \(\rho \) gives \(A_{ij}\rho A_{ij}^{\dagger }\ge 0\). Since the projections are Hermitian and sum to the identity, expansion and collection of the four double sums gives
Thus \(kT(\rho )-\rho \) is positive semidefinite.
13.3 Birkhoff’s theorem
Birkhoff’s theorem characterizes the convex geometry of doubly stochastic matrices.
The set of doubly stochastic matrices in \(M_d(\mathbb {R})\) is the convex hull of the \(d\times d\) permutation matrices, and its extreme points are exactly the permutation matrices. See [ Wol12 , Chapter 8, Theorem 8.6 ] . This statement concerns only doubly stochastic matrices; the corresponding assertion of Theorem 8.6 for doubly substochastic matrices is not included here.
Let \(\mathcal{DS}_d \subset M_d(\mathbb {R})\) be the set of \(d \times d\) doubly stochastic matrices and let \(\mathcal{P}_d\) be the set of \(d \times d\) permutation matrices. The Birkhoff–von Neumann decomposition gives
Indeed, every doubly stochastic matrix is a convex combination of permutation matrices, while convexity of \(\mathcal{DS}_d\) gives the converse inclusion in the first identity. Suppose that a permutation matrix \(P_\sigma \) is a convex combination \(P_\sigma =tA+(1-t)B\), where \(0{\lt}t{\lt}1\) and \(A,B\) are doubly stochastic. Nonnegativity and the row sums give
Thus \(A=B=P_\sigma \), so every permutation matrix is extreme. Conversely, a non-permutation matrix has a Birkhoff decomposition involving at least two distinct permutation matrices and is therefore not extreme, which proves the second identity.
13.4 Von Neumann entropy
For a Hermitian matrix \(\rho \in M_{D}(\mathbb {C})\) with eigenvalues \(\lambda _0,\ldots ,\lambda _{D-1}\), the von Neumann entropy is
where \(0\log 0:=0\).
If two Hermitian matrices are equal, then their von Neumann entropies are equal.
Substitute the equality of the matrices in the defining eigenvalue sum.
The zero matrix has zero von Neumann entropy: \(S(0)=0\).
All eigenvalues of the zero matrix are zero, and \(0\log 0=0\).
For any density matrix \(\rho \), \(S(\rho )\ge 0\).
Each eigenvalue \(\lambda _i\) of a density matrix satisfies \(0\le \lambda _i\le 1\), and \(-x\log x\ge 0\) on \([0,1]\).
The eigenvalues of a density matrix sum to \(1\).
Follows from \(\operatorname{tr}(\rho )=\sum _i\lambda _i=1\).
Each eigenvalue of a density matrix lies in \([0,1]\).
Non-negativity comes from positive semidefiniteness. The upper bound follows because the eigenvalues are non-negative and sum to \(1\).
For a density matrix \(\rho \in M_{D}(\mathbb {C})\) with \(D\ge 1\), one has \(S(\rho )\le \log D\).
By Jensen’s inequality applied to the concave function \(-x\log x\), the entropy is maximized when all eigenvalues equal \(1/D\).
For a density matrix \(\rho \) of rank \(r\), one has \(S(\rho )\le \log r\). This refines the dimension bound: only the nonzero eigenvalues contribute to the entropy, and there are exactly \(r\) of them.
The entropy is the \(-x\log x\) sum over all eigenvalues; the \(D-r\) zero eigenvalues contribute nothing. Jensen’s inequality applied to \(-x\log x\) over the \(r\) nonzero eigenvalues, with uniform weights \(1/r\), gives \(S(\rho )\le \log r\), the maximum attained when each nonzero eigenvalue equals \(1/r\). The rank equals the number of nonzero eigenvalues of the Hermitian matrix \(\rho \).
The von Neumann entropy of a Hermitian matrix is the \(-x\log x\) sum over the real parts of the roots of its characteristic polynomial \(\chi _\rho \):
The eigenvalues of a Hermitian matrix are exactly the roots of its characteristic polynomial, counted with multiplicity, so the eigenvalue sum defining \(S(\rho )\) equals the displayed sum over roots.
Let \(\rho \) be a Hermitian matrix and let \(\log \rho \) be its logarithm defined through the functional calculus. Then
A Hermitian matrix and its logarithm are simultaneously diagonalized by a unitary \(U\) with \(\rho =U\operatorname{diag}(\lambda _i)U^*\), so
Its negative is \(\sum _i({-}\lambda _i\log \lambda _i)=S(\rho )\). Zero eigenvalues contribute nothing under the convention \(0\log 0=0\), so no full-support assumption is required.
The logarithm is the totalized real logarithm, with \(\log x=\log \lvert x\rvert \) and \(\log 0=0\); both sides of the identity use it on every eigenvalue, so the equality holds for an arbitrary Hermitian matrix. It coincides with the physical entropy \(-\operatorname{tr}(\rho \log \rho )\) precisely when \(\rho \) is positive semidefinite. In that case (a density matrix \(\rho \), positive semidefinite with unit trace) the eigenvalues \(\lambda _i\) are non-negative, \(\rho \log \rho \) is Hermitian with real trace, and the real-part extraction is superfluous: the identity reduces to the standard expression \(S(\rho )=-\operatorname{tr}(\rho \log \rho )\).
For matrices \(\rho ,\sigma \in M_{D}(\mathbb {C})\), define the trace-log expression
On the physical domain where \(\rho \) is a density matrix and \(\sigma \) is positive definite, this is the Umegaki relative entropy.
For matrices \(\rho ,\sigma \in M_{D}(\mathbb {C})\),
Expand the matrix product over the difference \(\log \rho -\log \sigma \) and use linearity of the trace and of the real part.
For every matrix \(\rho \in M_{D}(\mathbb {C})\), one has \(D(\rho \Vert \rho )=0\).
The logarithmic difference \(\log \rho -\log \rho \) vanishes.
For every matrix \(\sigma \in M_{D}(\mathbb {C})\), one has \(D(0\Vert \sigma )=0\).
The trace-log expression is multiplied on the left by the zero matrix.
If \(\rho \) is Hermitian, then
Apply Lemma 13.4.13 to write
The trace-logarithm identity \(\operatorname{Re}\operatorname{tr}(\rho \log \rho )=-S(\rho )\) then gives the result.
Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices with \(\sigma \) of full rank. Then the relative entropy is non-negative, \(D(\rho \Vert \sigma )\ge 0\).
Diagonalize \(\rho =\sum _i p_i|e_i\rangle \! \langle e_i|\) and \(\sigma =\sum _j q_j|f_j\rangle \! \langle f_j|\) in their eigenbases, with all \(q_j{\gt}0\) since \(\sigma \) has full rank. The overlap numbers \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\) are non-negative with row sums and column sums equal to \(1\), because the two eigenbases are orthonormal. A trace computation gives
The row and column sums, together with the trace-one normalizations, give \(\sum _{i,j}P_{ij}(p_i-q_j)=\sum _i p_i-\sum _jq_j=0\). Hence
Each bracket is non-negative. If \(p_i{\gt}0\), put \(x=q_j/p_i{\gt}0\); the bracket is \(p_i(x-1-\log x)\geq 0\) by \(\log x\leq x-1\). If \(p_i=0\), the bracket equals \(q_j{\gt}0\). Therefore \(D(\rho \Vert \sigma )\geq 0\).
Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices satisfying the support condition \(\ker \sigma \subseteq \ker \rho \), that is, every vector annihilated by \(\sigma \) is annihilated by \(\rho \). Then the relative entropy is non-negative, \(D(\rho \Vert \sigma )\ge 0\).
Regularize \(\sigma \) by the trace-one perturbation \(\sigma _\varepsilon '=(1+\varepsilon D)^{-1}(\sigma +\varepsilon \mathbb {1})\), which is positive definite for every \(\varepsilon {\gt}0\), hence of full rank. The full-rank Klein inequality (Theorem 13.4.17) gives \(D(\rho \Vert \sigma _\varepsilon ')\ge 0\). The perturbation shares the eigenbasis \(\{ |f_j\rangle \} \) of \(\sigma \), so the cross term reduces to a scalar sum over the eigenvalues \(q_j\) of \(\sigma \),
For \(q_j{\gt}0\) the scalar logarithm converges to \(\log q_j\); for \(q_j=0\) the eigenvector \(|f_j\rangle \) lies in \(\ker \sigma \subseteq \ker \rho \), so its diagonal weight \(\langle f_j|\rho |f_j\rangle \) vanishes and the summand is identically zero. No eigenvalue-continuity input is needed, so \(D(\rho \Vert \sigma _\varepsilon ')\to D(\rho \Vert \sigma )\) as \(\varepsilon \to 0^+\) and the inequality passes to the limit.
Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices with \(\sigma \) of full rank. Then the relative entropy vanishes exactly when the states coincide, \(D(\rho \Vert \sigma )=0\iff \rho =\sigma \). Together with nonnegativity, this is the order property that makes \(D\) a divergence.
If \(\rho =\sigma \) then \(\log \rho -\log \sigma =0\) and \(D(\rho \Vert \sigma )=0\). For the converse, diagonalize \(\rho =\sum _i p_i|e_i\rangle \! \langle e_i|\) and \(\sigma =\sum _j q_j|f_j\rangle \! \langle f_j|\), with all \(q_j{\gt}0\) by full rank, and set \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\). As in Theorem 13.4.17,
a sum of non-negative terms: each is bounded below by the tangent inequality \(\log x\le x-1\) at \(x=q_j/p_i\), while the linear remainder telescopes through the doubly stochastic row and column sums, \(\sum _{i,j}P_{ij}(p_i-q_j)=\sum _ip_i-\sum _jq_j=0\). If the total vanishes then every term vanishes. On a row with \(p_i=0\) the term reads \(P_{ij}q_j\), so \(P_{ij}=0\); on a row with \(p_i{\gt}0\) the term forces the tangent inequality to be tight, \(\log (q_j/p_i)=q_j/p_i-1\), and strict concavity of the logarithm (\(\log x{\lt}x-1\) for \(x\neq 1\)) gives \(q_j=p_i\). Hence \(q_j=p_i\) whenever \(\langle e_i | f_j \rangle \neq 0\). Writing \(W=U_\rho ^\dagger U_\sigma \) for the overlap of the eigenvector unitaries, this matching says \(W\operatorname{diag}(q)=\operatorname{diag}(p)W\). Thus, conjugating \(\sigma \) into the eigenbasis of \(\rho \) and using the unitarity \(WW^\dagger =1\),
the spectral diagonal of \(\rho \). Therefore \(\sigma =U_\rho \operatorname{diag}(p)U_\rho ^\dagger =\rho \).
On pairs of positive definite matrices in \(M_{D}(\mathbb {C})\), the map \((\rho ,\sigma )\mapsto D(\rho \Vert \sigma )\) is jointly convex.
For \(s\in [0,1)\) consider the approximant
The trace \(\operatorname{tr}\rho \) is real-affine in the pair, hence convex, while \((\rho ,\sigma )\mapsto \operatorname{Re}\operatorname{tr}(\rho ^s\sigma ^{1-s})\) is jointly concave by the \(K=\mathbb {1}\) case of the Lieb concavity theorem (Corollary 7.7.14). Their difference, scaled by the non-negative factor \((1-s)^{-1}\), is therefore jointly convex, so each \(g_s\) is convex.
Writing \(\rho \) and \(\sigma \) in their eigenbases with eigenvalues \(p_i,q_j{\gt}0\) and overlap weights \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\), the approximant becomes the double sum \(g_s(\rho ,\sigma )=\sum _{i,j}(1-s)^{-1}(p_i-p_i^sq_j^{1-s})P_{ij}\). As \(s\to 1^-\) each per-pair term converges, via the scalar limit \((c^u-1)/u\to \log c\), to \(p_i(\log p_i-\log q_j)P_{ij}\), whose sum equals \(D(\rho \Vert \sigma )\). Thus \(g_s\to D\) pointwise on positive definite pairs, and since the pointwise limit of convex functions is convex, \(D\) is jointly convex.
On pairs of density matrices \((\rho ,\sigma )\) in \(M_{D}(\mathbb {C})\) satisfying the support condition \(\ker \sigma \subseteq \ker \rho \), the map \((\rho ,\sigma )\mapsto D(\rho \Vert \sigma )\) is jointly convex.
The domain is convex: for a strict convex combination of two such pairs the kernel of \(a\sigma _1+b\sigma _2\) is \(\ker \sigma _1\cap \ker \sigma _2\), since the quadratic forms of the positive semidefinite summands are non-negative and their positively weighted sum vanishes only when each does, and a positive semidefinite matrix annihilates exactly the vectors of zero quadratic form; the two pointwise support inclusions then give \(\ker \sigma _1\cap \ker \sigma _2\subseteq \ker \rho _1\cap \ker \rho _2\).
Regularize both arguments through the affine trace-one perturbation \(M_\varepsilon =(1+\varepsilon N)^{-1}(M+\varepsilon \mathbb {1})\), where \(N\) is the matrix dimension, which is positive definite for every \(\varepsilon {\gt}0\). Because the perturbation is affine, it commutes with the convex combination, so the four regularized endpoints and the regularized mixture all lie in the positive definite domain and the positive definite joint convexity (Theorem 13.4.20) gives the two-point inequality for the regularized pairs. As \(\varepsilon \to 0^+\), the regularized relative entropy converges on each pair of the support domain: \(D(\rho _\varepsilon \Vert \sigma _\varepsilon ) \to D(\rho \Vert \sigma )\). Indeed the perturbation shares the eigenbasis of its argument, so each trace-logarithm term is the diagonal sum \(\sum _jw_j(\varepsilon ) \log ((1+\varepsilon N)^{-1}(q_j+\varepsilon ))\), where \(q_j\) runs over the eigenvalues of \(\sigma \) and \(w_j(\varepsilon )\) is the corresponding diagonal weight. For \(q_j{\gt}0\), the scalar factor converges to \(\log q_j\), while at a zero eigenvalue the support condition makes the weight vanish, with \(w_j(\varepsilon )=(1+\varepsilon N)^{-1}\varepsilon \), so the summand is
by \(x\log x\to 0\) as \(x\to 0^+\). The two-point inequality therefore passes to the limit, giving joint convexity on the support domain.
For a Hermitian matrix \(A\), a real function \(f\), and a unitary \(U\), the continuous functional calculus satisfies
The conjugation \(\varphi :x\mapsto UxU^\dagger \) is a star-algebra automorphism of the matrix algebra, and the continuous functional calculus commutes with such automorphisms, \(\varphi (f(A))=f(\varphi (A))\); since \(\varphi (A)=UAU^\dagger \), this is the claim.
For a Hermitian matrix \(A\) and a unitary \(U\),
This is the \(f=\log \) case of Lemma 13.4.22: \(f(UAU^\dagger )=Uf(A)U^\dagger \) specialized to the real logarithm gives \(\log (UAU^\dagger )=U(\log A)U^\dagger \).
For Hermitian matrices \(\rho ,\sigma \) and a unitary \(U\),
Carry the two logarithms through the conjugation by Lemma 13.4.23, so that
where the middle equality is trace cyclicity together with \(U^\dagger U=1\).
For positive definite matrices \(\rho \) and \(\tau \),
The tensor product factors as \(\rho \otimes \tau =(\rho \otimes \mathbb {1})(\mathbb {1}\otimes \tau )\) into a pair of commuting positive definite matrices, so the logarithm of the product is the sum of the logarithms of the factors. Each unital embedding \(A\mapsto A\otimes \mathbb {1}\) and \(B\mapsto \mathbb {1}\otimes B\) is a continuous star-algebra homomorphism, hence commutes with the functional calculus, which carries each logarithm onto its factor: \(\log (\rho \otimes \mathbb {1})=\log \rho \otimes \mathbb {1}\) and \(\log (\mathbb {1}\otimes \tau )=\mathbb {1}\otimes \log \tau \).
For positive definite matrices \(\rho ,\sigma \) and a positive definite matrix \(\tau \) of unit trace,
Splitting each tensor logarithm by Lemma 13.4.25, the common \(\mathbb {1}\otimes \log \tau \) terms cancel in the difference, leaving \(\log (\rho \otimes \tau )-\log (\sigma \otimes \tau ) =(\log \rho -\log \sigma )\otimes \mathbb {1}\). Hence
where the trace of a tensor product factors as a product of traces. As \(\operatorname{tr}\tau =1\), the right-hand side is \(D(\rho \Vert \sigma )\).
Let \(\zeta \) be a primitive \(d\)-th root of unity and let \(i,j\) range over the residues modulo \(d\). Then
Each summand equals \(\xi ^b\) with \(\xi =\zeta ^i\overline{\zeta }^{\, j}\), and \(\xi ^d=1\). When \(i=j\) the base \(\xi \) equals \(1\) and the sum is \(d\). When \(i\ne j\) the base is a root of unity different from \(1\), so the geometric sum \((\xi -1)\sum _b\xi ^b=\xi ^d-1=0\) forces the sum to vanish; the equivalence \(\xi =1\iff i=j\) uses that \(\overline{\zeta }=\zeta ^{-1}\) and the injectivity of \(b\mapsto \zeta ^b\) on residues.
Fix a dimension \(d\ge 1\) and a primitive \(d\)-th root of unity \(\zeta \). The cyclic shift \(X\), the clock operator \(Z\), and every Weyl operator \(W(a,b)=X^aZ^b\) are unitary.
The cyclic shift is the permutation matrix of a cyclic permutation. The clock operator is diagonal, and each diagonal entry is a power of \(\zeta \), hence has modulus one. Thus \(X\) and \(Z\) are unitary, and so is every product \(X^aZ^b\).
Fix a dimension \(d\ge 1\) and a primitive \(d\)-th root of unity \(\zeta \), and let \(X\) be the cyclic shift \(|i\rangle \mapsto |i+1\rangle \) and \(Z=\operatorname{diag}(\zeta ^0,\ldots ,\zeta ^{d-1})\) the clock operator. For every matrix \(M\) on \(\mathbb {C}^d\), the uniform average of the conjugations by the \(d^2\) Weyl operators \(W(a,b)=X^aZ^b\) is the completely depolarizing channel:
The double average factors into a clock average followed by a shift average. The clock average \(\sum _bZ^bM(Z^b)^\dagger \) multiplies the entry \(M_{ij}\) by \(\sum _b\zeta ^{bi}\overline{\zeta ^{bj}}\), which by Lemma 13.4.27 is \(d\) when \(i=j\) and \(0\) otherwise; the result is \(d\) times the diagonal part of \(M\). The shift average \(\sum _aX^a(\operatorname{diag}v)(X^a)^\dagger \) cyclically permutes the diagonal entries, so each diagonal position receives the full sum \(\sum _kv_k=\operatorname{tr}M\), giving \((\operatorname{tr}M)\mathbb {1}\). Combining the two factors of \(d\) with the prefactor \(d^{-2}\) leaves \((\operatorname{tr}M/d)\mathbb {1}\).
Fix a dimension \(d_C\ge 1\) and a primitive \(d_C\)-th root of unity \(\zeta \). For every matrix \(M\) on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\), the uniform average of the conjugations by the \(d_C^2\) unitaries \(\mathbb {1}_S\otimes W(a,b)\) on the second factor is the partial trace over that factor tensored with the maximally mixed state \(\mathbb {1}_C/d_C\):
On each pair of blocks indexed by the first factor, the conjugation by \(\mathbb {1}_S\otimes W(a,b)\) acts as the Weyl conjugation \(W(a,b)(\cdot )W(a,b)^\dagger \) of the corresponding block of \(M\). Averaging over the \(d_C^2\) Weyl operators sends each block to the depolarizing channel by Theorem 13.4.29, replacing it by its trace times \(\mathbb {1}_C/d_C\). The block trace is exactly the corresponding entry of the partial trace \(\operatorname{tr}_C M\), so the average is \((\operatorname{tr}_C M)\otimes (\mathbb {1}_C/d_C)\).
For \(d_C\ge 1\), the matrix \(\tau _C=d_C^{-1}\mathbb {1}_C\) is positive definite and has trace one.
The identity is positive definite and \(d_C^{-1}{\gt}0\), so \(\tau _C\) is positive definite. Moreover, \(\operatorname{tr}\tau _C=d_C^{-1}\operatorname{tr}\mathbb {1}_C=1\).
Let \(A\) be a Hermitian matrix on a finite index set, and let \(e\) be a bijection onto another finite index set. Then
Reindexing is a star-algebra isomorphism and therefore commutes with the continuous functional calculus for the real logarithm.
For Hermitian matrices \(\rho ,\sigma \) on a finite index set and any bijection \(e\) from that set onto another finite set,
Write \(D(\rho \Vert \sigma ) =\operatorname{Re}\operatorname{tr}(\rho (\log \rho -\log \sigma ))\). The matrix logarithm is covariant under reindexing, \(\log (\rho _{e^{-1},e^{-1}}) =(\log \rho )_{e^{-1},e^{-1}}\), because reindexing is a star-algebra isomorphism and so commutes with the continuous functional calculus. The same isomorphism preserves products and the trace,
Taking \(M=\log \rho -\log \sigma \) and applying these three identities gives \(D(\rho _{e^{-1},e^{-1}}\Vert \sigma _{e^{-1},e^{-1}}) =D(\rho \Vert \sigma )\).
For positive definite matrices \(\rho ,\sigma \) on a tensor product of a system factor and an ancilla factor,
where \(\operatorname{tr}_C\) is the partial trace over the ancilla factor. This is the positive-definite base case; the source inequality on the support domain \(\ker \sigma \subseteq \ker \rho \) is Theorem 13.4.35.
By ancilla additivity (Theorem 13.4.26) the reduced-state relative entropy equals \(D\bigl((\operatorname{tr}_C\rho )\otimes (\mathbb {1}_C/d_C)\Vert (\operatorname{tr}_C\sigma )\otimes (\mathbb {1}_C/d_C)\bigr)\). By Lemma 13.4.30 each tensored reduced state is the uniform average of the conjugations by the \(d_C^2\) unitaries \(U_{ab}=\mathbb {1}_S\otimes W(a,b)\), so this is the relative entropy of a convex combination of the conjugated pairs \((U_{ab}\rho U_{ab}^\dagger ,U_{ab}\sigma U_{ab}^\dagger )\) with equal weights \(d_C^{-2}\). Joint convexity bounds it above by the same convex combination of the per-term relative entropies, each of which equals \(D(\rho \Vert \sigma )\) by unitary invariance (Theorem 13.4.24). As the weights sum to one, the bound is \(D(\rho \Vert \sigma )\).
For positive semidefinite \(\rho ,\sigma \) on a tensor product of a system factor and an ancilla factor of dimension \(d_C\), with the support condition \(\ker \sigma \subseteq \ker \rho \),
where \(\operatorname{tr}_C\) is the partial trace over the ancilla factor.
Let \(N\) be the dimension of the joint space. Regularize both arguments through the affine trace-shrinking perturbation \(M_\varepsilon =(1+N\varepsilon )^{-1}(M+\varepsilon \mathbb {1})\), which is positive definite for every \(\varepsilon {\gt}0\), so the positive-definite data-processing inequality (Theorem 13.4.34) gives
The right-hand side converges to \(D(\rho \Vert \sigma )\) as \(\varepsilon \to 0^+\), by the same shared-eigenbasis scalar-limit argument as the joint convexity on the support domain (Theorem 13.4.21).
For the left-hand side, the partial trace of the regularization is a differently scaled regularization of the marginal: because the partial trace is linear and \(\operatorname{tr}_C\mathbb {1}=d_C\mathbb {1}\),
This is the affine regularization of \(\operatorname{tr}_C M\) with the same scaling rate \(N\) but shift rate \(d_C\) and the smaller identity on the system factor. The support condition transfers to the marginals, \(\ker (\operatorname{tr}_C\sigma )\subseteq \ker (\operatorname{tr}_C\rho )\): a vector annihilated by \(\operatorname{tr}_C\sigma \) has vanishing marginal quadratic form, which splits into the non-negative joint quadratic forms of its single-ancilla lifts, so each lift lies in \(\ker \sigma \), hence in \(\ker \rho \), and reassembling the ancilla sum shows the vector lies in \(\ker (\operatorname{tr}_C\rho )\). With this support condition the arbitrary-rate affine regularization has the same both-arguments limit:
Passing the inequality through the two limits gives the support-domain bound.
Let \(\tau =\sum _i\lambda _i|i\rangle \! \langle i|\) be positive semidefinite. Its inverse square root on the support is
The inverse square root of a positive semidefinite matrix on its support is Hermitian.
The inverse square root of a positive semidefinite matrix on its support is positive semidefinite.
On the non-negative spectrum of \(\tau \), the defining function satisfies
The spectral functional calculus therefore gives \(f(\tau )\geq 0\).
If \(\tau \) is positive definite, then its inverse square root on the support is its ordinary inverse square root: \(\tau ^{-1/2}_{\mathrm{supp}}=(\sqrt\tau )^{-1}\).
Every eigenvalue of \(\tau \) is strictly positive. Hence the function defining the support inverse square root agrees on the spectrum with the reciprocal of the positive square-root function.
If \(P_\tau \) is the orthogonal projector onto the support of a positive semidefinite matrix \(\tau \), then
In an eigenbasis of \(\tau \), the left-hand side has eigenvalue zero when \(\lambda _i=0\) and eigenvalue \(\lambda _i^{-1/2}\lambda _i\lambda _i^{-1/2}=1\) otherwise.
If \(P_\tau \) is the orthogonal projector onto the support of a positive semidefinite matrix \(\tau \), then
The support inverse square root commutes with \(\tau \) by spectral functional calculus, so
where the last step is Lemma 13.4.40. Taking adjoints gives the identity with \(\tau \) on the left.
Let \(A\) and \(B\) be positive semidefinite, and let \(f\colon \mathbb {R}\to \mathbb {R}\) be multiplicative on the non-negative reals. Then
Diagonalize \(A\) and \(B\). In the resulting product eigenbasis, the eigenvalues of \(A\otimes B\) are \(a_i b_j\) with \(a_i,b_j\geq 0\), and the claim follows from \(f(a_i b_j)=f(a_i)f(b_j)\).
Let \(A\) and \(B\) be positive semidefinite. Their positive square roots, support inverse square roots, and support projections satisfy
Moreover,
The same support projection absorbs the square root:
If \(A\) is positive definite, then \(P_A=\mathbf1\).
Diagonalize both factors. The eigenvalues of \(A\otimes B\) are the products \(a_i b_j\). Both the square-root function and the function that equals \(x^{-1/2}\) for \(x{\gt}0\) and zero at \(x=0\) are multiplicative on the non-negative reals. For the support projectors, apply the sandwich identity to \(A\otimes B\), factor the support inverse square root and matrix products, and apply the sandwich identity to each factor. The cancellation identities follow entrywise. If \(A\) is positive definite, then \(A^{-1/2}_{\mathrm{supp}}=(\sqrt A)^{-1}\) and \(\sqrt A\) is invertible. Hence the cancellation identity gives \(P_A=\sqrt A(\sqrt A)^{-1}=\mathbf1\).
Let \(\tau \) be positive semidefinite and let \(c{\gt}0\). Then
For every finite-dimensional auxiliary space \(\mathcal H_R\),
The corresponding identity for an auxiliary left factor is
Consequently, for \(d_R{\gt}0\),
The scalar identity follows from the functional calculus and \(\sqrt{cx}=\sqrt c\sqrt x\) for \(c{\gt}0\). For the support-inverse function \(f(x)=x^{-1/2}\) when \(x{\gt}0\) and \(f(0)=0\), the functional calculus gives
This proves (61); the same argument with the identity as the left factor proves (62). Applying (60) with \(c=d_R^{-1}\) to \(\tau \otimes \mathbf1_R\) then gives (63).
Let \(\sigma \) be positive semidefinite on \(H_L\otimes H_R\), and set \(\tau =\operatorname{tr}_R\sigma \). The Petz transpose formula on the support of \(\tau \) is
This is the support formula of [ HJPW04 , Theorem 3, equation (8) ] . It is not asserted to be trace preserving on operators outside the support of \(\tau \).
For every matrix \(X\), the support Petz map is given by (64).
Let \(\rho _A\) and \(\rho _{BC}\) be positive semidefinite, and define
We use the canonical reassociation from \(A\times (B\times C)\) to \((A\times B)\times C\), so that the right partial trace removes \(C\).
Set \(\rho _B=\operatorname{tr}_C\rho _{BC}\), and let \(P_A\) and \(P_B\) be the support projections of \(\rho _A\) and \(\rho _B\). Then
Each identity is understood after the same canonical reassociation of the three tensor factors.
Expand the marginal in a product basis. The remaining identities follow from functional calculus for positive semidefinite tensor products. The support projection is obtained by multiplying the factorized support inverse square root on both sides of the marginal.
If \(P_A\) is the support projection of \(\rho _A\), define
The operator-Schmidt rank below is a generic bipartite-matrix fact, with no tensor-network content; it is relocated here from the MPDO area-law chapter, whose diagonal-cut and pure-state support-compression bounds cite it across the chapter boundary.
For a bipartite density operator \(\rho \in M_{d_A}(\mathbb {C})\otimes M_{d_B}(\mathbb {C})\), its operator-Schmidt rank is the least integer \(r\) for which there are matrices \(A_t\in M_{d_A}(\mathbb {C})\) and \(B_t\in M_{d_B}(\mathbb {C})\) satisfying
No Hermiticity or positivity condition is imposed on the factors. This is the definition in [ DlCDN19 , Equation (1) ] ; the same formula defines the rank of an arbitrary complex bipartite matrix.
For the reference in (65), the raw Petz map for \(\operatorname{tr}_C\) satisfies
Equivalently, for every product operator \(A_0\otimes X_B\),
The formulas use the canonical identification \((H_A\otimes H_B)\otimes H_C\cong H_A\otimes (H_B\otimes H_C)\).
This is the globally valid ambient-space form of [ HJPW04 , equation (10) ] . The literal identity-tensored formula in that equation is obtained on the support of \(\rho _A\), or after choosing an extension away from that support. For singular \(\rho _A\), the raw map on the full matrix algebra contains the compression \(X\mapsto P_AXP_A\).
Substitute the tensor factorizations of the square root and marginal support inverse into the Petz sandwich. The first-factor terms reduce to \(P_AA_0P_A\). This proves the formula for product operators. A finite product-operator decomposition proves the linear-map identity.
Let \(X\) be an operator on \(H_A\otimes H_B\). If
then
If \(\rho _A\) is positive definite, then \(P_A=\mathbf1_A\), and this identity holds for every \(X\).
When \(\rho _A\) is singular, no global identity-tensored formula is asserted for the raw map outside the displayed support. Nor is the generic trace-preserving completion in Definition 13.4.63 asserted to factor: its complementary projection is \(\mathbf1_{AB}-P_A\otimes P_B\), which need not be the identity on \(A\) tensored with a projection on \(B\).
The support condition makes \(\mathcal S_{P_A}\) act as the identity on \(X\). If \(\rho _A\) is positive definite, its support projection is the identity, so the condition holds on the full matrix algebra.
Let \(d_A{\gt}0\) and let \(\rho _{BC}\) be positive semidefinite. Define
Under the canonical reassociation from \(A\times (B\times C)\) to \((A\times B)\times C\), the right partial trace removes \(C\).
For the reference in (71),
Expand the two partial traces in a product basis. The sum over the \(d_A\) diagonal entries cancels the factor \(d_A^{-1}\). Trace invariance under a partial trace then gives the last identity.
If \(\rho _{BC}\) is positive semidefinite and \(\rho _B=\operatorname{tr}_C\rho _{BC}\), then
The marginal identity gives \(\sigma _{AB}=d_A^{-1}\mathbf1_A\otimes \rho _B\). Apply the support-inverse scaling identity and the tensor identity for an auxiliary left factor. Since \(d_A{\gt}0\), \((\sqrt{d_A^{-1}})^{-1}=\sqrt{d_A}\).
For the reference in (71), the support Petz map for \(\operatorname{tr}_C\) factors as
after the canonical identification \((H_A\otimes H_B)\otimes H_C\cong H_A\otimes (H_B\otimes H_C)\). This is the maximally mixed specialization of [ HJPW04 , equation (10) ] . It is the tensor-product identity for the raw Petz support formula, not the Hayashi–Koashi–Imoto block decomposition.
Functional calculus gives \(\sqrt{\sigma _{ABC}} =d_A^{-1/2}\mathbf1_A\otimes \sqrt{\rho _{BC}}\). Theorem 13.4.55 gives the corresponding factorization of the marginal support inverse. The scalar factors cancel in the Petz sandwich. The identity follows first for \(A_0\otimes X_B\) and then for every operator by a finite sum of product operators.
If \(P_B\) is the support projector of \(\rho _B\), then the support projector of \(d_A^{-1}\mathbf1_A\otimes \rho _B\) is
Insert the factorized support inverse into \(P_{AB}=\sigma _{AB,\mathrm{supp}}^{-1/2} \sigma _{AB}\sigma _{AB,\mathrm{supp}}^{-1/2}\). The scalar factors cancel, and the remaining sandwich is \(\mathbf1_A\otimes (\rho _{B,\mathrm{supp}}^{-1/2}\rho _B \rho _{B,\mathrm{supp}}^{-1/2})=\mathbf1_A\otimes P_B\).
The support Petz transpose map \(\mathcal R_\sigma \) is completely positive.
For an orthonormal basis \((e_r)_r\) of \(H_R\), let \(J_r:H_L\to H_L\otimes H_R\) be given by \(J_r(v)=v\otimes e_r\). Then
Thus the support Petz map has a rectangular Kraus representation.
Let \(P_\tau \) be the orthogonal projector onto the support of \(\tau =\operatorname{tr}_R\sigma \). Then, for every matrix \(X\),
Cyclicity of the trace and the defining property of the partial trace give
By (55), this equals \(\operatorname{tr}(P_\tau X)\).
Put \(Q_\tau =\mathbf1_L-P_\tau \) and \(\omega _R=\operatorname{tr}_L\sigma \). The complementary term is
The complementary preparation term \(\mathcal C_\sigma \) is completely positive.
The map \(X\mapsto Q_\tau XQ_\tau \) has the single Kraus operator \(Q_\tau \). Adjoining the positive semidefinite matrix \(\omega _R\) is completely positive, and the composition of these two maps is completely positive.
If \(\sigma \) has trace one, then
Since \(\operatorname{tr}(\omega _R)=\operatorname{tr}(\sigma )=1\) and \(Q_\tau ^2=Q_\tau \), factorization and cyclicity of the trace give \(\operatorname{tr}(\mathcal C_\sigma (X)) =\operatorname{tr}(Q_\tau XQ_\tau ) =\operatorname{tr}(Q_\tau ^2X) =\operatorname{tr}(Q_\tau X)\).
The completed Petz map is
If \(\sigma \) is positive semidefinite with trace one, then \(\widehat{\mathcal R}_\sigma \) is completely positive and trace preserving.
For the reference in (71), the chosen complementary preparation term factors as
after the canonical reassociation of the three tensor factors. This follows from the chosen support completion in (77), not from [ HJPW04 , equation (10) ] .
From (75), \(Q_{AB}=\mathbf1_A\otimes Q_B\), while \(\operatorname{tr}_{AB}\sigma _{ABC}=\rho _C\). Hence
A finite product-operator decomposition gives the result for every input.
For the reference in (71),
after canonical reassociation of the three tensor factors. The raw support-map summand is the maximally mixed specialization of [ HJPW04 , equation (10) ] . The complementary summand comes from the chosen support completion in (77). This theorem does not assert a Hayashi–Koashi–Imoto decomposition.
If \(\rho _{BC}\) is positive semidefinite with trace one, then \(\widehat{\mathcal R}_{\sigma _{ABC}}\) is completely positive and trace preserving.
By (72), \(\operatorname{tr}(\sigma _{ABC})=\operatorname{tr}(\rho _{BC})=1\). Apply the channel property of the completed Petz map.
If \(P_\tau XP_\tau =X\), then \(\widehat{\mathcal R}_\sigma (X)=\mathcal R_\sigma (X)\).
The assumption implies \(Q_\tau XQ_\tau =0\). Hence \(\mathcal C_\sigma (X)=0\), and adding the complementary term to \(\mathcal R_\sigma (X)\) does not change its value.
Let \(A\) be Hermitian, with spectral decomposition \(A=U\operatorname{diag}(\lambda _i)U^\dagger \). Its support projection is
The logarithm identity below is a generic support-projection fact, with no tensor-network content; it is relocated here from the MPDO area-law chapter, whose tripartite strong-subadditivity argument cites it across the chapter boundary.
Let \(A\) and \(B\) be positive semidefinite, with \(P_A\) and \(P_B\) the orthogonal projections onto their respective ranges. Then
Here the logarithm is extended by zero on the kernel. Neither matrix is required to be positive definite, and either index set may be empty. This auxiliary identity is project-derived rather than a theorem of CPSV16.
Diagonalize \(A\) and \(B\). On a product eigenvector with eigenvalues \(a,b\geq 0\), equation (83) becomes
If \(a,b{\gt}0\), this is the ordinary product identity for the logarithm. If either eigenvalue is zero, both sides vanish. Conjugating the resulting diagonal identity by the product eigenbasis proves the formula.
Let \(A\) and \(B\) be Hermitian matrices such that \(\ker A\subseteq \ker B\). If \(P_A\) is the support projection of \(A\), then \(P_A B P_A=B\).
The complementary projection \(\mathbf1-P_A\) has range contained in \(\ker A\), and hence in \(\ker B\). Thus \(B(\mathbf1-P_A)=0\). Taking adjoints gives \((\mathbf1-P_A)B=0\), so \(B=P_A B\), and therefore \(P_A B P_A=P_A B=B\).
For \(w\in H_L\) and a distinguished basis vector \(e_r\in H_R\), define the lift \(J_r w=w\otimes e_r\in H_L\otimes H_R\).
Let \(X\) be a matrix on \(H_L\otimes H_R\). For basis indices \(i,s\) and \(r\),
Since \((J_r w)_{(j,c)}=w_j\mathbf1_{c=r}\), expansion of the matrix-vector product gives
Let \(X\) be a matrix on \(H_L\otimes H_R\), let \(w\in H_L\), and let \((e_r)_r\) be the distinguished orthonormal basis of \(H_R\). Then
Expanding the matrix products and the partial trace gives \(\sum _{i,j,r}\overline{w_i}X_{(i,r),(j,r)}w_j\) on both sides.
Let \(\sigma \) be positive semidefinite on \(H_L\otimes H_R\), and let \(\rho \) be any matrix on the same space such that \(\ker \sigma \subseteq \ker \rho \). Then
If \(w\in \ker (\operatorname{tr}_R\sigma )\), then
Each summand is non-negative and therefore vanishes. Positive semidefiniteness shows that \(\sigma (w\otimes e_r)=0\) for every \(r\), so the joint kernel inclusion gives \(\rho (w\otimes e_r)=0\). Summing the diagonal components over \(r\) yields \((\operatorname{tr}_R\rho )w=0\).
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(H_L\otimes H_R\) such that \(\ker \sigma \subseteq \ker \rho \). Then
Kernel inclusion descends through the partial trace, giving \(\ker (\operatorname{tr}_R\sigma )\subseteq \ker (\operatorname{tr}_R\rho )\). Theorem 13.4.71 therefore yields \(P_\tau (\operatorname{tr}_R\rho )P_\tau =\operatorname{tr}_R\rho \), where \(\tau =\operatorname{tr}_R\sigma \). By Theorem 13.4.68,
For every positive semidefinite \(\sigma \),
By (55), the middle factor reduces to \(P_\tau \otimes \mathbf1_R\). Marginal-support absorption gives \((\mathbf1_{LR}-P_\tau \otimes \mathbf1_R)\sigma =0\), so
because that matrix is positive semidefinite of trace zero. Therefore \(\sqrt\sigma (P_\tau \otimes \mathbf1_R)\sqrt\sigma =\sigma \).
For every positive semidefinite \(\sigma \),
The marginal satisfies \(P_\tau \tau P_\tau =\tau \). Therefore the completed channel agrees with the support Petz map at \(\tau \), and Theorem 13.4.77 gives \(\widehat{\mathcal R}_\sigma (\tau )=\sigma \).
Let \(\rho \) be a Hermitian matrix indexed by a finite set \(J\), and let \(e : I \to J\) be a bijection from a finite set \(I\). The reindexed matrix on \(I\) with entries \(\rho _{e(i)\, e(j)}\) has the same von Neumann entropy as \(\rho \).
By Lemma 13.4.9 the entropy depends only on the characteristic polynomial. Reindexing conjugates \(\rho \) by a permutation matrix, so \(\chi _{(\rho _{e(i)\, e(j)})}=\chi _\rho \), and the entropies coincide.
Let \(A \in M_{m \times n}(\mathbb {C})\) and \(B \in M_{n \times m}(\mathbb {C})\). Then the charpoly-root entropy sum is invariant under the cyclic swap \(AB \mapsto BA\):
with roots counted with algebraic multiplicity.
The rectangular characteristic-polynomial identity gives \(X^n\chi _{AB}=X^m\chi _{BA}\). Thus the two characteristic polynomials have the same nonzero roots, with multiplicity; the additional zero roots contribute nothing because \(0\log 0=0\).
For matrices \(A \in M_{m \times n}(\mathbb {C})\) and \(B \in M_{n \times m}(\mathbb {C})\) such that \(AB\) and \(BA\) are Hermitian, \(S(AB)=S(BA)\).
For density matrices \(\omega \) and \(\tau \), \(S(\omega \otimes \tau )=S(\omega )+S(\tau )\).
The tensor product is unitarily conjugate to the diagonal of eigenvalue products \(\lambda _i\mu _j\). Using
and the unit eigenvalue sums collapses the double sum to \(S(\omega )+S(\tau )\).
For a density matrix \(\omega \) and a scalar \(c\), \(S(c\, \omega )=c\, S(\omega )-c\log c\).
The eigenvalues of \(c\, \omega \) are \(c\lambda _i\), so the entropy is \(\sum _i-(c\lambda _i)\log (c\lambda _i)\). The splitting identity
together with the unit eigenvalue sum \(\sum _i\lambda _i=1\) gives
the stated form.
For a family of Hermitian matrices \(M_j\), the block-diagonal direct sum satisfies
Each block diagonalizes by a unitary congruence, and the block-diagonal assembly of the block unitaries diagonalizes the direct sum, whose eigenvalue multiset is the disjoint union of the block eigenvalue multisets. Hence
Let \(A=\sum _j M_j\) be a finite sum of Hermitian matrices. Suppose there are operators \(P_j\) such that
Then \(S(A)=\sum _j S(M_j)\). The operators \(P_j\) need only resolve the support of \(A\); their sum need not be the identity on the ambient space. This is the support form of the direct-sum entropy identity used in [ CPGSV16 , Appendix C.2, lines 1760–1770 ] .
Write \(A=XY\), where \(X\) maps the direct sum of the labelled ambient spaces to the original space by the matrices \(M_j\), and \(Y\) maps back by the operators \(P_j\). The support and annihilation identities give \(YX=\bigoplus _jM_j\). Reversing the two rectangular factors preserves the nonzero eigenvalues, while any additional eigenvalues are zero. Entropy is therefore unchanged, and additivity on the block diagonal gives the result.
Let \(\omega _j\) be density matrices and let \(p_j\geq 0\). Then
In particular, when the \(p_j\) form a probability distribution, this is
This is the entropy identity used in [ CPGSV16 , Appendix C.2, lines 1760–1770 ] .
Entropy is additive over the orthogonal blocks. Applying the scaled-state formula to each block gives \(S(p_j\omega _j)=-p_j\log p_j+p_jS(\omega _j)\), and summing over \(j\) gives the result.
Suppose \(p_j{\gt}0\) and \(L_j\leq R_j\) for every \(j\). If \(\sum _jp_jL_j=\sum _jp_jR_j\), then \(L_j=R_j\) for every \(j\). This is the positivity argument applied to strong subadditivity in [ CPGSV16 , Appendix C.2, lines 1770–1780 ] .
Each number \(p_j(R_j-L_j)\) is non-negative, and their sum vanishes. Therefore every one vanishes. Since \(p_j{\gt}0\), it follows that \(R_j-L_j=0\).
13.5 Tripartite partial traces
For a tripartite matrix \(\rho _{ABC}\) on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B} \otimes \mathbb {C}^{d_C}\), the partial trace over \(A\) is
The partial trace over \(C\) is
The partial trace over \(A\) and \(C\) is
If \(\rho _{ABC}\) is Hermitian, then \(\operatorname{tr}_A(\rho _{ABC})\), \(\operatorname{tr}_C(\rho _{ABC})\), and \(\operatorname{tr}_{AC}(\rho _{ABC})\) are all Hermitian. The same holds for bipartite partial traces \(\operatorname{tr}_A(\rho _{AB})\) and \(\operatorname{tr}_B(\rho _{AB})\).
Follows from \(\overline{\rho _{ji}} = \rho _{ij}\) applied entry-wise inside the summation defining each partial trace.
13.6 Strong subadditivity
Let \(\rho \) be a density matrix on \(A \otimes R\) with reduced state \(\rho _R = \operatorname{tr}_A \rho \), and suppose the support condition \(\ker ((\mathbb {1}_A / d_A) \otimes \rho _R) \subseteq \ker \rho \) holds. Then
When \(\rho _R\) is singular the tensor logarithm of the reference does not split. Regularize the reduced state by the affine perturbation \(\rho _{R,\varepsilon } = (1 + d_R\varepsilon )^{-1}(\rho _R + \varepsilon \mathbb {1})\), positive definite for \(\varepsilon {\gt} 0\). The reference \((\mathbb {1}_A / d_A) \otimes \rho _{R,\varepsilon }\) is then a positive definite tensor product, whose logarithm splits, so the cross trace term evaluates to \(-\log d_A + \operatorname{Re}\operatorname{tr}(\rho _R\log \rho _{R,\varepsilon })\). As \(\varepsilon \to 0^+\) the support condition makes the zero eigenvalues of \(\rho _R\) contribute nothing, so
Since \(D(\rho \| \sigma ) = -S(\rho ) - \operatorname{Re}\operatorname{tr}(\rho \log \sigma )\), the evaluation follows.
Let \(\rho \) be a positive semidefinite operator on \(A \otimes R\) with reduced state \(\rho _R = \operatorname{tr}_A \rho \). Then the singular reference \((\mathbb {1}_A / d_A) \otimes \rho _R\) satisfies the support condition \(\ker ((\mathbb {1}_A / d_A) \otimes \rho _R) \subseteq \ker \rho \).
A vector \(v\) annihilated by \((\mathbb {1}_A / d_A) \otimes \rho _R\) is annihilated by \(\mathbb {1}_A \otimes \rho _R\), since \(\mathbb {1}_A / d_A\) is invertible. Let \(P\) be the orthogonal projection onto the range of \(\rho _R\). The complementary lift \(\mathbb {1}_A \otimes (\mathbb {1}- P)\) then fixes \(v\), while it annihilates \(\rho \) on the left because the reduced state of \(\rho \) on \(R\) is supported on the range of \(P\). Hence \(\rho v = 0\).
For every tripartite matrix \(\rho _{ABC}\),
For indices \((a_1,b_1)\) and \((a_2,b_2)\), the corresponding matrix entry is
This is the corresponding entry of \((\mathbb {1}_A/d_A)\otimes \rho _B\).
For a tripartite density matrix \(\rho _{ABC}\),
Read the inequality as one instance of data processing under the partial trace over \(C\), with the singular reference state \(\sigma _{ABC} = (\mathbb {1}_A / d_A) \otimes \rho _{BC}\). Being a density operator is not enough to place the pair \((\rho _{ABC}, \sigma _{ABC})\) in the relative-entropy domain, which is the kernel inclusion \(\ker \sigma _{ABC} \subseteq \ker \rho _{ABC}\); the marginal support lemma supplies it. Against this reference the relative entropy of each pair evaluates to an entropy difference:
For a singular reference the tensor logarithm does not split, so each evaluation regularizes the reduced state through the affine perturbation \(\rho _{R,\varepsilon } = (1 + d_R\varepsilon )^{-1}(\rho _R + \varepsilon \mathbb {1})\), which is positive definite, and passes to the limit
where the zero eigenvalues of \(\rho _R\) contribute nothing because the kernel inclusion makes the corresponding diagonal weights vanish. Data processing on the singular support domain under the partial trace over \(C\), which sends \(\rho _{ABC} \mapsto \rho _{AB}\) and \((\mathbb {1}_A / d_A) \otimes \rho _{BC} \mapsto (\mathbb {1}_A / d_A) \otimes \rho _B\), reads
Substituting (100) and (101) into (102) cancels the common \(\log d_A\) and rearranges to the claimed inequality.
For a positive definite tripartite density matrix \(\rho _{ABC}\),
This is the positive definite case of Theorem 13.6.4.
Read the inequality as one instance of data processing under the partial trace over \(C\), with reference state \(\sigma _{ABC} = (\mathbb {1}_A / d_A) \otimes \rho _{BC}\), which is positive definite because \(\rho _{BC}\) is. Against this reference the relative entropy of each pair evaluates to an entropy difference:
Each equality follows from the tensor logarithm split and the adjoint of the partial trace. Data processing under the partial trace over \(C\), which sends \(\rho _{ABC} \mapsto \rho _{AB}\) and \((\mathbb {1}_A / d_A) \otimes \rho _{BC} \mapsto (\mathbb {1}_A / d_A) \otimes \rho _B\), reads
Substituting (103) and (104) into (105) cancels the common \(\log d_A\) and rearranges to the claimed inequality.
A tripartite density matrix \(\rho _{ABC}\) satisfies SSA equality if
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices such that \(\ker \sigma \subseteq \ker \rho \), and let \(\tau \) be positive semidefinite with \(\operatorname{tr}\tau =1\). Then
Write \(P_\rho \), \(P_\sigma \), and \(P_\tau \) for the support projections. The kernel inclusion gives \(P_\sigma \rho P_\sigma =\rho \) and \(\rho P_\sigma =\rho \), while \(\rho P_\rho =\rho \) and \(\tau P_\tau =\tau \). Substituting the singular tensor-logarithm formulas into the relative entropy and using these four identities cancels the two contributions containing \(\log \tau \). Factoring the trace of the remaining tensor product gives
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\) such that \(\ker \sigma \subseteq \ker \rho \), and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For a primitive \(d_C\)-th root of unity, put \(U_{ab}=\mathbb {1}_S\otimes W(a,b)\) and
Then, for every \(a,b\),
This is a scalar equality-propagation prerequisite for [ HJPW04 , Theorem 3 and equation (8) ] ; it neither characterizes equality in joint convexity nor asserts recovery.
Let \(\tau _C=d_C^{-1}\mathbb {1}_C\). The twirl identity gives \(\overline X=(\operatorname{tr}_C X)\otimes \tau _C\). The support inclusion passes to the partial traces:
Hence support-domain ancilla additivity and the saturation hypothesis give
Every \(U_{ab}\) is unitary, and unitary invariance gives
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\) such that \(\ker \sigma \subseteq \ker \rho \), and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For a primitive \(d_C\)-th root of unity, put \(U_{ce}=\mathbb {1}_S\otimes W(c,e)\) and
Then
This scalar identity is associated with the finite Jensen step in the Weyl proof of data processing. It is a prerequisite for [ HJPW04 , Theorem 3 and equation (8) ] ; it neither characterizes equality in joint convexity nor asserts recovery.
Unitary invariance gives, for every \(c,e\),
Theorem 13.6.8 gives \(D(\overline\rho \Vert \overline\sigma )=D(\rho \Vert \sigma )\). Substituting into the right-hand side gives
If \(A\) and \(B\) are positive definite and \(c{\gt}0\), then
The functional-calculus identity \(\log (cA)=(\log c)\mathbb {1}+\log A\), and its analog for \(B\), show that the scalar logarithmic terms cancel in \(\log (cA)-\log (cB)\). Linearity of the trace then gives the result.
Let \(A\) and \(B\) be positive semidefinite and suppose that \(\ker B\subseteq \ker A\). If \(c{\gt}0\), then
On the support of a positive semidefinite matrix,
The support projection \(P_A\) fixes \(A\) on the right. The kernel inclusion also makes \(P_B\) fix \(A\) on the right. Hence the two terms containing \(\log c\) cancel in the trace-log difference, and linearity of the trace gives the result.
Let \(\rho \) and \(\sigma \) be positive definite, and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For the uniformly weighted Weyl conjugates \(A_g=d_C^{-2}U_g\rho U_g^\dagger \) and \(B_g=d_C^{-2}U_g\sigma U_g^\dagger \), put \(A=\sum _gA_g\) and \(B=\sum _gB_g\). Then
Let \(a,b{\gt}0\). Then the function
is integrable on \((0,\infty )\), and
This is the scalar normalization \((\mathrm{intspec})\) in Jenčová–Ruskai, arXiv:0903.2895v4, §2.1, lines 406–413.
The integrand is the derivative of \(a(\log (1+t)-\log (a+tb))\). Its value at \(t=0\) is \(-a\log a\), whereas its limit as \(t\to \infty \) is \(-a\log b\). The derivative has constant sign, according as \(a-b\) is positive or negative, and is therefore integrable. The fundamental theorem of calculus gives the stated value.
Let \(A\) and \(B\) be Hermitian matrices of the same size, with spectral resolutions
Write \(w_{ij}=\langle u_i,v_j\rangle \), and use the total real logarithm: \(\log 0=0\), while \(\log x=\log |x|\) for \(x{\lt}0\). Then
This is an algebraic totalized extension to arbitrary Hermitian matrices of the homogeneous trace-log identity \((\mathrm{J1})\), which Jenčová–Ruskai state for strictly positive matrices in arXiv:0903.2895v4, lines 277–287.
Expand both trace terms in eigenbases. Unitarity of the overlap matrix gives \(\sum _j|w_{ij}|^2=1\), so
Subtraction gives (117).
Let \(A\) and \(B\) be positive semidefinite matrices of the same size and suppose that \(\ker B\subseteq \ker A\). With the spectral notation of Theorem 13.6.14, define, for \(t{\gt}0\),
Thus no ordinary quotient with \(\alpha _i=\beta _j=0\) is used. Define also
The function \(I_{A,B}\) is integrable on \((0,\infty )\), and
The trace-log identity \((\mathrm{J1})\), the scalar normalization \((\mathrm{intspec})\), and its matrix form \((\mathrm{intAB})\) occur in Jenčová–Ruskai, arXiv:0903.2895v4, at lines 277–287, 406–413, and 423–427, respectively. The support-domain extension is given at lines 717–720.
This theorem concerns the spectral expression \(I_{A,B}\). The next theorem identifies it with the coordinate-free left-right quadratic form underlying the finite Weyl formula.
If \(\beta _j=0\), the kernel inclusion gives \(\alpha _i|w_{ij}|^2=0\), and both \(r_{ij}\) and \(e_{ij}\) are defined to be zero. If \(\beta _j{\gt}0\), the scalar integral applies when \(\alpha _i{\gt}0\), while the term is identically zero when \(\alpha _i=0\). Since the double sum is finite, integration term by term gives
The kernel inclusion also shows that replacing each \(\beta _j=0\) summand in Theorem 13.6.14 by zero does not change its value. Hence that theorem identifies the sum in (123) with \(D(A\Vert B)\).
Let \(A\) and \(B\) be positive semidefinite matrices of the same size, let \(P_B\) be the orthogonal projection onto the support of \(B\), and, for \(t{\gt}0\), set
Write \(S_t^+\) for the generalized inverse that vanishes on \(\ker S_t\). Define
With the spectral notation of Theorem 13.6.15,
Consequently,
The function in (126) is continuous on \((0,\infty )\). If \(\ker B\subseteq \ker A\), then \(A P_B=A\), so \(Q_{A P_B}(t)\) equals the quadratic form with source \(\operatorname{vec}(A^{\top })\).
This is the support-projected form of \((\mathrm{intAB})\) in Jenčová–Ruskai, arXiv:0903.2895v4, §2.1, lines 423–427, with the support convention at lines 717–720.
Diagonalize \(A\) and \(B\) and write \(W=U_A^\ast U_B\). In these coordinates, the two equations
have entries
For every positive semidefinite \(S\), the identities \(S^+S=SS^+=P_S\) imply that \(Sx=b\) gives \(\langle b,S^+b\rangle =\langle b,x\rangle \). Applying this identity to (127) and (128), with (131) and (134), gives (124) and (125). Their coefficientwise combination is the scalar function appearing in Lemma 13.6.13. Continuity follows term by term from the finite sums.
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size. For \(t{\gt}0\), write
Then
No kernel inclusion between \(A_i\) and \(B_i\) is required. This lemma is a positive-semidefinite support-domain extension of the positive-definite calculation in equations \((\mathrm{Mj})\), \((\mathrm{eq:Schz1})\), and \((\mathrm{eq:Schwzt})\) at lines 1313–1343 of Jenčová–Ruskai, arXiv:0903.2895v4; their generalized-inverse notation is given at lines 254–262. The paper does not state this extension. Its later singular equality theorem at lines 761–785 assumes \(\ker B_i\subseteq \ker A_i\) and is not asserted here. The subsequent singular entropy-equality passage is recorded in the TNLean paper-gap note [ con26p ] .
Put
The source equation for the support relative-modular operator shows that \(b_i\) lies in the support of \(S_i\). The same argument applied to \(\sum _i A_i\) and \(\sum _i B_i\) shows that \(\sum _i b_i\) lies in the support of \(\sum _i S_i\). The support-resolvent residual identity writes the displayed defect as a sum of nonnegative quadratic residuals.
Under the hypotheses and notation of Lemma 13.6.17, suppose that the source-\(B\) defect vanishes:
Put
and let \(P_{S_i}\) be the support projection of \(S_i\). Then
This is the fixed-parameter support-domain residual step behind \((\mathrm{basiceq})\) in Jenčová–Ruskai, arXiv:0903.2895v4, lines 652–660 and 788–790. No kernel inclusion is needed for this algebraic implication.
The source vectors lie in the supports of their left–right operators, and their sum lies in the support of \(S_\Sigma \). The vanishing real defect and the support-resolvent residual identity force every quadratic residual to vanish. The common-solution conclusion follows after projection to the support of each \(S_i\).
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size satisfying \(\ker B_i\subseteq \ker A_i\). For \(t{\gt}0\), write
Then
The kernel inclusions for the summands imply \(\ker (\sum _i B_i)\subseteq \ker (\sum _i A_i)\). This lemma is the source-\(A\) support-domain extension of the positive-definite residual calculation in equations \((\mathrm{Mj})\), \((\mathrm{eq:Schz1})\), and \((\mathrm{eq:Schwzt})\) at lines 1313–1343 of Jenčová–Ruskai, arXiv:0903.2895v4. The paper does not state this fixed-parameter extension separately. The lemma does not assert that equality of relative entropies makes the defect vanish.
Put
The inclusion \(\ker B_i\subseteq \ker A_i\) gives \(A_iP_{B_i}=A_i\). The source-\(A\) left–right equation therefore places \(a_i\) in the support of \(S_i\). Positivity shows that a vector annihilated by \(\sum _i B_i\) is annihilated by every \(B_i\), and hence by every \(A_i\). Thus the summed source also lies in the support of the summed operator. The support-resolvent residual identity writes the displayed difference as a sum of nonnegative quadratic residuals.
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size satisfying \(\ker B_i\subseteq \ker A_i\). Put
Then the function
is continuous on \((0,\infty )\) and integrable there, and
This is the finite-family support-domain form of \((\mathrm{intspec})\) and \((\mathrm{intAB})\) in Jenčová–Ruskai, arXiv:0903.2895v4, lines 406–431, with the positive-semidefinite convention and kernel hypotheses at lines 717–720 and 766–785.
The kernel inclusions for the summands imply \(\ker B_\Sigma \subseteq \ker A_\Sigma \). Apply the one-pair support-domain integral representation to every \((A_i,B_i)\) and to \((A_\Sigma ,B_\Sigma )\). The pointwise left–right identities identify the difference of the spectral integrands with \(F\); linearity of trace cancels the terms involving \(\operatorname {tr}(B_i)\). Finite summation commutes with the integral. Continuity follows from the corresponding one-pair continuity theorem and finite summation.
Under the hypotheses and notation of Theorem 13.6.20, for every \(t{\gt}0\) one has
This is the coefficient-correct combination of the two source defects in \((\mathrm{intAB})\) of Jenčová–Ruskai, arXiv:0903.2895v4, lines 423–435.
Both defects are nonnegative. Since \(t{\gt}0\) and \(1+t{\gt}0\), multiplying the source-\(B\) defect by \(t\), adding the source-\(A\) defect, and dividing by \(1+t\) preserves nonnegativity.
Under the hypotheses of Theorem 13.6.20, suppose
Then, for every \(t{\gt}0\),
This is the source-\(B\) pointwise-vanishing passage in Jenčová–Ruskai, arXiv:0903.2895v4, lines 433–435, 652–674, and 788–790. The common projected resolvent conclusion is stated downstream in Theorem 13.6.35.
The relative-entropy equality and (135) make the integral of the nonnegative function \(F\) vanish. Hence \(F=0\) almost everywhere. Its continuity on the open positive half-line upgrades this to \(F(t)=0\) for every \(t{\gt}0\). Both defects are nonnegative, so \(\operatorname {Def}_A(t)+t\operatorname {Def}_B(t)=0\), and \(t{\gt}0\) forces \(\operatorname {Def}_B(t)=0\).
Let \(A\) and \(B\) be positive definite, with spectral resolutions \(A=\sum _i\alpha _i|u_i\rangle \! \langle u_i|\) and \(B=\sum _j\beta _j|v_j\rangle \! \langle v_j|\). Put \(w_{ij}=\langle u_i,v_j\rangle \). Let \(L_A\) and \(R_B\) denote left and right multiplication, \(L_A(X)=AX\) and \(R_B(X)=XB\). Then, for \(t{\gt}0\),
Both quadratic forms are continuous on \((0,\infty )\). Moreover,
and the integrand in (138) is continuous and integrable on the positive half-line. This is the positive-definite spectral route of Jenčová–Ruskai, arXiv:0903.2895v4, §4.
Vectorization sends \(L_A+tR_B\) to \(A\otimes \mathbb {1}+t\mathbb {1}\otimes B^{\mathsf T}\). The vectors \(u_i\otimes \overline{v_j}\) diagonalize this matrix with eigenvalues \(\alpha _i+t\beta _j\), which proves (136) and (137). Insert these identities into (138), use \(\sum _i|w_{ij}|^2=1\), and apply Lemma 13.6.13 term by term. The same finite spectral sum proves continuity and integrability.
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(S_i\) be a positive-semidefinite matrix, let \(P_i\) be its support projection, and let \(b_i\) lie in its support. Write
Assume also that \(b\) lies in the support of \(S\), and put \(x=Gb\). Then
In particular,
Jenčová and Ruskai give the positive-definite residual expansion in equations \((\mathrm{Mj})\) and \((\mathrm{eq:Schz1})\) of the Appendix to arXiv:0903.2895v4. The support assumptions make the same expansion valid for the generalized inverses of the \(S_i\) and of \(S\).
Since \(G_iS_i=S_iG_i=P_i\) and \(P_i b_i=b_i\), expansion of the \(i\)th summand gives
Sum over \(i\). The support assumption on \(b\) gives \(Sx=SGb=b\), so the last three terms combine to \(-\langle b,Gb\rangle \).
Under the hypotheses and notation of Lemma 13.6.24, suppose that
Then, for every \(i\in I\),
This is the support-domain form of the common-resolvent equation \((\mathrm{basiceq})\) in Section 3.1 of Jenčová–Ruskai, arXiv:0903.2895v4. Its residual calculation is the one in the Appendix, equations \((\mathrm{Mj})\) and \((\mathrm{eq:Schz1})\).
Lemma 13.6.24 writes the complex defect in (140) as a finite sum of non-negative quadratic forms. Its imaginary part therefore vanishes automatically, while the hypothesis makes its real part vanish. Hence \(G_i(b_i-S_ix)=0\) for every \(i\). Multiplication by \(S_i\) shows that \(P_i(b_i-S_ix)=0\). Both \(b_i\) and \(S_ix\) lie in the support of \(S_i\), so \(b_i=S_ix\). Multiplication by \(G_i\) now gives (141).
Let \(\rho \) and \(\sigma \) be positive definite on \(\mathcal H_S\otimes \mathbb C^{d_C}\), let \(q=d_C^{-2}\), and put
where \(U_g=\mathbf1_S\otimes W_g\). For \(t{\gt}0\), let
Then
This is the positive-definite, fixed-\(t\) identity in Jenčová–Ruskai, arXiv:0903.2895v4, Appendix, lines 1313–1343. It does not infer zero defect from equality of relative entropies and makes no assertion about singular supports.
Write \(S_g\) for the positive definite matrix representing \(T_g\) under the vectorization \(X\mapsto \operatorname{vec}(X^{\mathsf T})\), and put \(b_g=\operatorname{vec}(B_g^{\mathsf T})\) and \(x=S^{-1}\sum _g b_g\). For the residual \(r_g=b_g-S_gx\), direct expansion gives
Since \(S_g^{1/2}\) is invertible,
Summing proves the identity.
Under the notation of Theorem 13.6.26, set \(\Gamma _t=T^{-1}(A)\). Then, for every \(t{\gt}0\),
This is the second defect family in Jenčová–Ruskai, arXiv:0903.2895v4, §4 and Appendix.
Apply the residual identity of Theorem 13.6.26 with \(a_g=\operatorname{vec}(A_g^{\mathsf T})\) and \(\bar a=\operatorname{vec}(A^{\mathsf T})\). Thus
which is the asserted source-\(A\) identity.
For the positive definite finite-Weyl family, let
and define \(\operatorname{Def}_B(t)\) analogously, with the real parts of the corresponding source-\(B\) pairings. Then both defects are non-negative and continuous for \(t{\gt}0\), and
The integrand is continuous, non-negative, and integrable on \((0,\infty )\). The coefficient of the source-\(B\) defect is exactly \(t\), as prescribed by the integral formula and equality analysis of Jenčová–Ruskai, arXiv:0903.2895v4, §4.
Apply Theorem 13.6.23 to each pair \((A_g,B_g)\) and to \((A,B)\), and interchange the finite sum with the integral. The trace terms cancel because \(B=\sum _gB_g\). The two fixed-resolvent identities write the defects as sums of squared norms, proving nonnegativity. Continuity and integrability follow from the corresponding spectral assertions before taking the finite difference.
Under the hypotheses and notation of the preceding theorem, suppose that
Then \(T_g^{-1}(B_g)=T^{-1}(B)\) for every Weyl index \(g\). In particular this holds for the identity Weyl element \(g=(0,0)\). This is the common-resolvent conclusion in Jenčová–Ruskai, arXiv:0903.2895v4, §4, lines 652–674; its squared-defect input is in Appendix, lines 1313–1343. This fixed-\(t\), positive-definite conclusion does not assert that equality of relative entropies implies the scalar hypothesis in (142).
The hypothesis (142) is a finite sum of squared norms. Each term is non-negative, so every term vanishes. Thus \(S_g^{-1/2}(b_g-S_gx)=0\) for every \(g\). Invertibility of \(S_g^{-1/2}\) gives \(b_g=S_gx\), and hence \(S_g^{-1}b_g=x=S^{-1}\sum _hb_h\).
Under the positive-definite finite-Weyl hypotheses, suppose that \(\sum _gD(A_g\Vert B_g)-D(A\Vert B)=0\). Then \(\operatorname{Def}_B(t)=0\) for every \(t{\gt}0\), and consequently \(T_g^{-1}(B_g)=T^{-1}(B)\) for every \(t{\gt}0\) and every Weyl index \(g\). In particular, equality of relative entropy under the right partial trace implies this conclusion by Theorem 13.6.12. This is the positive-definite conclusion of the equality argument in Jenčová–Ruskai, arXiv:0903.2895v4, §4 and Appendix. It makes no assertion at \(t=0\) or for singular inputs.
The integrand in Theorem 13.6.28 is non-negative and has integral zero, hence it vanishes almost everywhere. Its continuity improves this to vanishing at every \(t{\gt}0\). Since \(\operatorname{Def}_A(t)\geq 0\), \(\operatorname{Def}_B(t)\geq 0\), and \(t/(1+t){\gt}0\), it follows that \(\operatorname{Def}_B(t)=0\). Theorem 13.6.29 now gives the common solution.
Let \(S,T\) be positive semidefinite matrices and let \(x\) be a vector. If \((t\mathbf1+S)^{-1}x=(t\mathbf1+T)^{-1}x\) for every \(t{\gt}0\), then \(\sqrt S\, x=\sqrt T\, x\). More generally, for any fixed matrix \(Q\), if \(Q(t\mathbf1+S)^{-1}x=Q(t\mathbf1+T)^{-1}x\) for every \(t{\gt}0\), then \(Q\sqrt S\, x=Q\sqrt T\, x\).
In particular, for positive definite \(A,B\) and every \(t{\gt}0\), put \(\Delta _{A,B}=A\otimes (B^{-1})^{\mathsf T}\). The source-\(B\) left–right resolvent satisfies
and
These are the positive-square-root specializations of the passage from relative modular resolvents to analytic functions of the relative modular operator in Jenčová–Ruskai, arXiv:0903.2895v4, lines 658–680.
Use the Löwner integral representation of the power \(p=1/2\). Its integrand at \(t{\gt}0\) is \(f_t(S)=t^{-1/2}\mathbf1-t^{1/2}(t\mathbf1+S)^{-1}\). Applied to \(x\), this is \(f_t(S)x=t^{-1/2}x-t^{1/2}(t\mathbf1+S)^{-1}x\). The hypothesis therefore gives \(Q(f_t(S)x)=Q(f_t(T)x)\) for every \(t{\gt}0\). Since \(M\mapsto Q(Mx)\) is a bounded linear map from \(M_{n}(\mathbb {C})\) to \(\mathbb C^n\),
and similarly for \(T\). Integration therefore gives \(Q\sqrt S\, x=Q\sqrt T\, x\).
For the second assertion, factor \(L_A+tR_B=(\Delta _{A,B}+t\mathbf1)R_B\). Applying the inverse to \(B=R_B(\mathbf1)\) gives the shifted relative modular resolvent on \(\mathbf1\). Finally,
as follows from uniqueness of the positive square root.
For matrices \(M\) and \(P\), column-stacking vectorization satisfies
This is the Kronecker vectorization identity \((B\otimes A)\operatorname{vec}(X)=\operatorname{vec}(AXB^{\mathsf T})\) with \(A=P^{\mathsf T}\), \(B=\mathbf1\), and \(X=M^{\mathsf T}\).
Let \(A,B\) be positive semidefinite, let \(t{\gt}0\), set \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\), and let \(P_B\) be the support projection of \(B\). Then
Put \(C=\mathbf1\otimes B^{\mathsf T}\). The support generalized-inverse identity gives
Together with \(B^{\mathsf T}P_B^{\mathsf T}=B^{\mathsf T}\), this factors the left two operators as
The shifted relative-modular matrix is positive definite and hence invertible, so
The Kronecker vectorization identity gives \(C\operatorname{vec}(\mathbf1^{\mathsf T})=\operatorname{vec}(B^{\mathsf T})\), which is the claim.
Let \(A,B\) be positive semidefinite, let \(t{\gt}0\), set
and let \(P_B\) be the support projection of \(B\). Then
Consequently,
This is the one-pair algebraic identification used in the singular equality argument of Jenčová–Ruskai, arXiv:0903.2895v4, lines 783–790. It does not assert that equality of relative entropies gives a common resolvent for a finite family.
Put
The generalized-inverse identities give \(SP=CR\), \(CD=P\), and \(P^2=P\). The preceding lemma shows that \(y=PR^{-1}\operatorname{vec}(\mathbf1^{\mathsf T})\) satisfies \(Sy=\operatorname{vec}(B^{\mathsf T})\), and \(Py=y\). Every vector \(v\) satisfying \(Pv=v\) lies in the range of \(S\), since
In particular, \(y\) lies in the range of \(S\), so the support projection \(P_S\) of \(S\) satisfies \(P_Sy=y\). Therefore
Pairing this vector equality with \(\operatorname{vec}(B^{\mathsf T})\) and taking real parts gives the quadratic identity.
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices satisfying \(\ker B_i\subseteq \ker A_i\). Put
and suppose that
Then, for every \(i\in I\) and \(t{\gt}0\),
Thus the shifted relative-modular resolvents agree on \((\ker B_i)^\perp \). The local support projection is essential; no ambient equality is asserted. This is the support-restricted conclusion of Jenčová–Ruskai, arXiv:0903.2895v4, lines 766–793.
Relative-entropy equality makes the source-\(B\) defect vanish for every positive parameter. The common left–right solution theorem and the one-pair projected relative-modular identity then compare the local solution with the summed solution. The local right-support projection absorbs both the support projection of the local left–right operator and the right-support projection of the summed reference matrix. Applying it to the common-solution identity gives the displayed equality.
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices satisfying \(\ker \sigma \subseteq \ker \rho \), and suppose that
Fix a primitive \(d_C\)-th root of unity, put \(U_g=\mathbf1\otimes W_g\), and define the unweighted Weyl family
Then, for every Weyl index \(g\) and every \(t{\gt}0\),
In particular, the zero Weyl index gives the projected resolvent of the original pair \((\rho ,\sigma )\). No ambient resolvent equality is asserted. The support convention is that of Jenčová–Ruskai, arXiv:0903.2895v4, lines 717–720, and their common-resolvent equality argument is at lines 766–793. This is one analytic step toward the recovery implication of Hayden–Jozsa–Petz–Winter, Theorem 3 and equation (8); it is not the direct-sum Markov decomposition invoked in CPSV16, Lemma Lsigma3.
Unitary conjugation preserves positive semidefiniteness and transports the kernel inclusion to every pair \((A_g,B_g)\). The kernel of the sum of the \(B_g\) is contained in the kernel of the sum of the \(A_g\). The finite Weyl Jensen equality and Lemma 13.6.11 give
Theorem 13.6.35 applied to this finite family gives the displayed projected resolvent.
Let \(X\) be a matrix on \(\mathcal H_S\otimes \mathbb C^{d_C}\) and let \(U_g=\mathbf1\otimes W_g\) for the \(d_C^2\) Weyl indices. Then
The uniform Weyl twirl is \((\operatorname{tr}_CX)\otimes d_C^{-1}\mathbf1_C\). Multiplying its defining identity by \(d_C^2\) gives the unweighted sum.
Let \(A\) and \(B\) be positive semidefinite, and write \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\). Then
The support inverse square root is positive semidefinite, and
Uniqueness of the positive square root gives the corresponding Kronecker factorization. Applying the Kronecker vectorization identity to \(\operatorname{vec}(\mathbf1^{\mathsf T})\) gives the stated equation.
Let \(A,B,C,D\) be positive semidefinite matrices. Write \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\) and \(D^+=(D^{-1/2}_{\mathrm{supp}})^2\), and let \(P_B\) be the orthogonal projection onto \((\ker B)^\perp \). If, for every \(t{\gt}0\),
where the inverses act as relative-modular superoperators on matrices, then
This is the square-root specialization of the support functional-calculus passage in Jenčová–Ruskai, arXiv:0903.2895v4, lines 788–793. The projection \(P_B\) is essential: the source gives the common generalized resolvents only after restriction to \((\ker B)^\perp \). Deriving this restricted equality requires the preceding singular equality argument and its kernel hypotheses; these are supplied for finite families by Theorem 13.6.35.
By Lemma 13.6.38,
Setting \(Q=\mathbf1\otimes P_B^{\mathsf T}\) and applying Lemma 13.6.31 gives
Since \(\operatorname{vec}(X^{\mathsf T})=\operatorname{vec}(Y^{\mathsf T})\) implies \(X=Y\), injectivity of vectorization gives the conclusion.
Let \(\rho \) and \(\sigma \) be positive definite and suppose that \(D(\rho \Vert \sigma ) =D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). Put
Then
Hence the raw partial-trace Petz map satisfies \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \). This is the positive-definite case only; no singular-support conclusion is asserted.
Apply the common-resolvent theorem to the identity Weyl summand and the Weyl average. Lemma 13.6.31 converts the result to equality of the two square-root ratios. For \(c=d_C^{-2}\), the uniform scalar in the identity summand cancels through
The Weyl twirl identifies the average with the displayed maximally mixed extensions.
Taking the adjoint product of the ratio equality yields
Multiplication by \(\sqrt\sigma \) on both sides gives the sandwich. The raw Petz recovery identity follows from Theorem 13.6.42.
Let \(\tau \) be positive semidefinite on \(\mathcal H_S\), let \(X\) be a matrix on \(\mathcal H_S\), and let \(\overline\tau =\tau \otimes d_C^{-1}\mathbf1_C\). Then
By Lemma 13.4.44, each outer factor contributes \(\sqrt{d_C}\). The scalar coefficient cancels because \(\sqrt{d_C}\, d_C^{-1}\sqrt{d_C} =d_C^{-1}(\sqrt{d_C})^2=1\). The remaining matrix product is the asserted unital tensor embedding.
Let \(\rho \) and \(\sigma \) be matrices on \(\mathcal H_S\otimes \mathbb C^{d_C}\), with \(\sigma \) positive semidefinite, and set \(\overline\sigma =(\operatorname{tr}_C\sigma )\otimes d_C^{-1}\mathbf1_C\). If the identity summand obeys the support sandwich identity
then the raw support Petz map recovers \(\rho \): \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \). This is the algebraic reduction in [ HJPW04 , Theorem 3, equation (8) ] . It does not derive the support sandwich identity from equality of relative entropies.
Lemma 13.6.41 rewrites the middle three factors as
The support Petz formula therefore identifies the left-hand side of the assumed sandwich identity with \(\mathcal R_\sigma (\operatorname{tr}_C\rho )\).
In the finite Weyl coordinates \(\mathbb C^{d_S}\otimes \mathbb C^{\mathbb Z/d_C\mathbb Z}\), let \(\rho \) and \(\sigma \) be positive semidefinite matrices satisfying \(\ker \sigma \subseteq \ker \rho \), and suppose that
Put
If \(P_\sigma \) is the orthogonal projection onto \((\ker \sigma )^\perp \), then
Consequently, \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \) for the raw support Petz map. The projected relative-modular argument follows Jenčová–Ruskai, arXiv:0903.2895v4, lines 766–793, and the recovery formula is [ HJPW04 , Theorem 3, equation (8) ] . This result does not assert the middle-space direct-sum decomposition in [ CPGSV16 , Lemma Lsigma3, lines 1351–1363 ] .
Use the zero Weyl index in Theorem 13.6.36. Lemma 13.6.39 gives the projected square-root ratio for \((\rho ,\sigma )\) and the unweighted Weyl sums. Lemma 13.6.37 identifies those sums with \(d_C^2\overline\rho \) and \(d_C^2\overline\sigma \). The factors \(\sqrt{d_C^2}\) and \((\sqrt{d_C^2})^{-1}\) cancel, giving the first displayed equality.
Taking the adjoint product of that equality gives
The support projection absorbs \(\sqrt\sigma \) on both sides, while \(\ker \sigma \subseteq \ker \rho \) gives \(P_\sigma \rho P_\sigma =\rho \). Multiplying the last equality by \(\sqrt\sigma \) on the left and right therefore yields the support sandwich. Theorem 13.6.42 gives the raw Petz recovery identity.
Let \(e_L:H_L\to H_L'\) and \(e_R:H_R\to H_R'\) be bijections, and write \(Z^{e_L\otimes e_R}\) for the corresponding simultaneous relabelling of the rows and columns of \(Z\). Then
The first identity follows by changing the summation index in the partial trace. Functional calculus commutes with a relabelling by a bijection, so both \(\sqrt{\sigma }\) and the support inverse square root of \(\operatorname{tr}_R\sigma \) have the same covariance. The embedding \(X\mapsto X\otimes \mathbf1_R\) and matrix multiplication also commute with the product relabelling, which gives (144).
Let \(\rho \) and \(\sigma \) be density operators on the finite-dimensional product \(H_L\otimes H_R\) such that \(\ker \sigma \subseteq \ker \rho \). If
then the completed partial-trace Petz channel associated with \(\sigma \) recovers \(\rho \):
This is the right-partial-trace forward implication of [ HJPW04 , Theorem 3, equation (8) ] . The additional off-support term belongs to the trace-preserving completion and vanishes on \(\operatorname{tr}_R\rho \); it is not part of the formula in the cited equation.
Choose bijections from \(H_L\) to a finite standard basis and from \(H_R\) to a finite cyclic basis. Relative entropy, the support inclusion, and the equality (145) are unchanged under this product relabelling. Theorem 13.6.43 gives raw Petz recovery in the cyclic coordinates, and Lemma 13.6.44 transports the identity back to \(H_L\otimes H_R\). Finally, Theorem 13.4.76 identifies the completed channel with the raw support formula at \(\operatorname{tr}_R\rho \), giving (146).
For every positive semidefinite operator \(\omega _{XY}\),
No invertibility assumption is made on either marginal. This is [ HJPW04 , Equation (4) ] .
The lifted support projections \(P_X\otimes \mathbf1_Y\) and \(\mathbf1_X\otimes P_Y\) both fix \(\omega _{XY}\). Substituting Theorem 13.4.70 into the cross term of the relative entropy and using the two partial-trace adjoint identities gives
The asserted formula follows from \(D(\rho \, \| \sigma )=-S(\rho )-\operatorname{Re}\operatorname{tr}(\rho \log \sigma )\).
Let \(\omega _{XY}\) be positive semidefinite. Then
Equivalently, \(\omega _{XY}\) is supported on \(\operatorname {supp}\omega _X\otimes \operatorname {supp}\omega _Y\).
The support projection of \(\omega _X\otimes \omega _Y\) is \(P_X\otimes P_Y\). The lifted projections \(P_X\otimes \mathbf1_Y\) and \(\mathbf1_X\otimes P_Y\) both fix \(\omega _{XY}\), so their product \(P_X\otimes P_Y\) fixes it as well. Every vector annihilated by the product of the marginals is therefore annihilated by \(\omega _{XY}\).
Let \(\rho _{ABC}\) be a tripartite density matrix. Then equality holds in strong subadditivity if and only if relative-entropy data processing under the partial trace over \(C\) is saturated for the pair \(\rho _{ABC}\) and \(\rho _A\otimes \rho _{BC}\):
Moreover, \(\operatorname{tr}_C(\rho _A\otimes \rho _{BC})=\rho _A\otimes \rho _B\). This is the product-marginal formulation in [ HJPW04 , Equations (5)–(7) ] .
The two relative entropies are
respectively. Cancelling the common term \(S(\rho _A)\) shows that their equality is precisely
Entrywise, the partial-trace identity is
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity. The raw Petz support map for the reference \(\rho _A\otimes \rho _{BC}\) recovers \(\rho _{ABC}\) from \(\rho _{AB}\). Moreover,
where \(\widehat{\mathcal R}_{\rho _{BC}}\) is a trace-preserving completely positive extension of the support formula. This is equation (11) of [ HJPW04 ] , obtained from Theorem 3 and equation (8) together with the supported form of equation (10).
No marginal is required to be invertible. If \(\rho _A\) is singular, the displayed factorization is asserted on the supported input \(\rho _{AB}\) only; it is not a global factorization of an ambient product-reference completion.
Equality in strong subadditivity gives equality in relative-entropy data processing for \(\rho _{ABC}\) and \(\rho _A\otimes \rho _{BC}\). Lemma 13.6.47 supplies the required support inclusion, and Theorem 13.6.45 gives raw recovery after the canonical reassociation of the tensor factors.
The state \(\rho _{AB}\) is fixed on both sides by \(P_A\otimes \mathbf1_B\), so the supported product-reference factorization removes the first-factor support compression. It is also fixed by \(\mathbf1_A\otimes P_B\); hence every \(B\)-block lies in the support of \(\rho _B\). The completed local channel therefore agrees with its raw support formula on every block, proving both displayed identities.
For a finite-dimensional system \(A\), there is a finite family of effects \((M_s)_s\), with \(0\leq M_s\leq \mathbf1_A\), whose complex linear span is the full matrix algebra on \(A\). One member is the identity. For an operator \(X\) on \(A\otimes B\), define its conditional slice by
The family may be chosen from the four rank-one effects occurring in the polarization identity, scaled so that every member is bounded by the identity.
If two operators \(X,Y\) on \(A\otimes B\) have
for every member of the finite effect family, then \(X=Y\).
The polarization identity expresses every matrix unit as a complex linear combination of the selected rank-one effects. Hence any linear map vanishing on all selected effects vanishes on the full matrix algebra. Apply this to each matrix entry of the difference of the two conditional slice maps.
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and put \(\rho _{AB}=\operatorname{tr}_C\rho _{ABC}\). If \(\widehat{\mathcal R}:B\to B\otimes C\) is the completed recovery channel, define
Then \(\varphi \) is trace-preserving and completely positive and
For each selected effect, let
The indices with \(p_s\neq 0\) form a finite nonempty set. For each such index, positivity of \(\xi _s\) makes \(p_s\) real and positive, and
is a density operator and \(\varphi (\mu _s)=\mu _s\). Consequently a Kraus representation of \(\varphi \) gives a single preserving operation for this finite nonempty density family.
This is the conditional-family construction in [ HJPW04 , Theorem 6, lines 493–505 ] . Effects with \(p_s=0\) are omitted before normalization; no artificial conditional state is introduced for them.
Compose equation (11) with the partial trace over \(C\) to obtain the fixed point identity for \(\rho _{AB}\). Conditional slicing commutes with every linear map on \(B\), so \(\varphi (\xi _s)=\xi _s\). Positivity of \(\rho _{AB}\) and \(M_s\) gives positivity of \(\xi _s\). On the nonzero-trace subtype, scaling by \(p_s^{-1}\) preserves positivity and makes the trace one; the fixed-point equation is preserved by the same scaling. The identity effect has probability \(\operatorname{tr}(\rho _{AB})=1\), proving nonemptiness.
Let \(\rho _{ABC}\) be a tripartite density matrix and set \(\sigma _{ABC}=(\mathbf1_A/d_A)\otimes \rho _{BC}\). Then equality holds in strong subadditivity if and only if relative-entropy data processing under the partial trace over \(C\) is saturated for this pair:
By Lemma 13.6.3, the reference on the right is \((\mathbf1_A/d_A)\otimes \rho _B\).
This is an equivalent hypothesis-free criterion with a different reference state. The exact product-marginal formulation is Theorem 13.6.48.
The two relative entropies are
respectively. The marginal support lemma places both singular references in the relative-entropy domain. Cancelling the common \(\log d_A\) shows that equality of the two relative entropies is precisely
Let \(\rho \) be a Hermitian operator on \(A\otimes B\otimes C\), and let \(e_A:A'\to A\), \(e_B:B'\to B\), and \(e_C:C'\to C\) be bijections. Define \(\rho '\) by
If \(\rho \) satisfies equality in strong subadditivity, then so does \(\rho '\).
The four operators entering the equality are related by the induced bijections:
These identities follow by changing variables in the three finite sums defining the partial traces. Entropy invariance under reindexing makes the corresponding four entropy terms equal, so the strong-subadditivity equality for \(\rho \) gives the equality for \(\rho '\).
A Hayashi Markov decomposition of a tripartite state \(\rho _{ABC}\) consists of a finite direct-sum decomposition
together with a unitary change of basis on \(B\), a probability vector \((p_j)_j\), and density matrices \(\rho _{A B_j^L}\) and \(\rho _{B_j^R C}\) such that, in the adapted basis, the state becomes
The terminology follows Hayashi’s presentation of quantum Markov structure [ Hay06 ] ; the block decomposition used by the MPDO argument is the structure theorem of [ HJPW04 ] .
The finite-dimensional operator-algebraic step of the Hayden–Jozsa–Petz–Winter derivation of the Koashi–Imoto theorem ( [ HJPW04 , Appendix A ] ) begins with the common invariant algebra of a family of jointly invariant states.
Let \(\rho _1,\ldots ,\rho _K\) be density matrices in \(M_{D}(\mathbb {C})\), \(K\ge 1\). A trace-preserving completely positive Kraus family \(F\) preserves \(\rho _1,\ldots ,\rho _K\) if \(F\rho _k=\rho _k\) for every \(k\). Write
for the set of such operations – non-empty since the identity operation belongs to it – and
for their common average.
Every \(F\in \mathbf F\) satisfies \(F\bar\rho =\bar\rho \).
Linearity turns \(F\rho _k=\rho _k\) for every \(k\) into \(F\bar\rho =\frac1K\sum _kF\rho _k=\frac1K\sum _k\rho _k=\bar\rho \).
Assume in addition that the common average \(\bar\rho \) of Definition 13.6.56 is positive definite. Hayden–Jozsa–Petz–Winter instead reduce to this case by shrinking to the joint support of \(\rho _1,\ldots ,\rho _K\); that reduction is not re-derived here. By Lemma 13.6.57, \(\bar\rho \) is a positive definite fixed point of every \(F\in \mathbf F\), so Theorem 10.2.10 makes the fixed-point set of each adjoint map,
a \(*\)-subalgebra of \(M_{D}(\mathbb {C})\). The common invariant algebra is
and \(X\in A_0\) if and only if \(F^*(X)=X\) for every \(F\in \mathbf F\).
Under the hypotheses of Definition 13.6.58, there are \(L\in \mathbb {N}\), positive dimensions \(d_0,\ldots ,d_{L-1}\) and multiplicities \(m_0,\ldots ,m_{L-1}\) with \(\sum _\ell d_\ell m_\ell =D\), and a unitary \(U\in M_{D}(\mathbb {C})\) such that a matrix \(A\in M_{D}(\mathbb {C})\) belongs to \(A_0\) exactly when
for some matrices \(B_\ell \in M_{d_\ell }(\mathbb {C})\).
Direct specialization of Theorem 10.6.9 to \(S=A_0\).
Under the hypotheses of Definition 13.6.58, there are finitely many operations \(F_1,\ldots ,F_M\in \mathbf F\), \(M\ge 1\), with
Because \(M_{D}(\mathbb {C})\) is finite-dimensional, among the finite intersections \(\bigcap _{F\in S}A_F\) over finite sets \(S\subseteq \mathbf F\) containing a fixed operation, one of minimal dimension already equals \(A_0\): enlarging \(S\) by one more operation either leaves the intersection unchanged or strictly decreases its dimension, and dimension cannot decrease indefinitely.
Under the hypotheses of Definition 13.6.58, there is \(F_0\in \mathbf F\) with \(A_0=A_{F_0}\).
Take \(F_1,\ldots ,F_M\in \mathbf F\) as in Theorem 13.6.60 and pool their Kraus operators, each rescaled by \(1/\sqrt M\), into a single Kraus family representing
Rescaling every pooled Kraus operator by \(1/\sqrt M\) turns the sum of the \(M\) trace-preservation identities into a single one, so \(F_0\in \mathbf F\); rescaling a Kraus operator by a nonzero scalar does not change the operators it commutes with, so \(A_{F_0}=A_{F_1}\cap \cdots \cap A_{F_M}=A_0\).
Fix \(F_0\in \mathbf F\) with \(A_0=A_{F_0}\) as in Theorem 13.6.61. Let \(P_0\) be the finite-dimensional mean-ergodic projection of the Schrödinger map \(F_0\), and define \(P_0^*\) as its trace-pairing adjoint.
For every \(X\in M_{D}(\mathbb {C})\),
The trace pairing transports every iterate of \(F_0^*\) to the corresponding iterate of \(F_0\), and hence transports each finite Cesàro average. Apply Theorem 10.1.5 to \(F_0\) and use nondegeneracy of the trace pairing.
The map \(P_0^*\) is positive, unital, and idempotent, and its range is exactly \(A_0\).
Theorem 10.1.9 applied to \(F_0\) gives that \(P_0^*\) is positive, unital, and idempotent, with range equal to the fixed-point set of \(F_0^*\), which is \(A_{F_0}=A_0\).
Under the hypotheses of Definition 13.6.58, there are \(L\in \mathbb {N}\), positive dimensions \(d_\ell ,m_\ell \), a unitary \(U\), and density matrices \(\sigma _\ell \in M_{m_\ell }(\mathbb {C})\) such that, for every member \(\rho _k\) of the invariant family, there are matrices \(X_{\ell ,k}\in M_{d_\ell }(\mathbb {C})\) satisfying
The density matrices \(\sigma _\ell \) and the decomposition are common to the whole family.
No positivity or trace normalization is asserted for the matrices \(X_{\ell ,k}\). Thus this is the full-support fixed-point block form underlying HJPW Appendix A, lines 853–856, not the normalized decomposition of Property 1.
Apply Theorem 10.8.16 to the Schrödinger map of the single operation \(F_0\) from Theorem 13.6.61. Its adjoint satisfies the Schwarz inequality by the Kraus form and trace preservation, and it fixes the positive-definite common average. Since \(F_0\rho _k=\rho _k\) for every \(k\), each family member belongs to the one fixed-point space \(U(\bigoplus _\ell \sigma _\ell \otimes M_{d_\ell }(\mathbb {C}))U^\dagger \).
Suppose in addition that every \(\rho _k\) is positive semidefinite with trace one. Under the explicit positive-definite common-average hypothesis of Definition 13.6.58, the common decomposition of Theorem 13.6.65 admits numbers \(q_{\ell |k}\geq 0\) and density matrices \(\tau _{\ell |k}\in M_{d_\ell }(\mathbb {C})\) such that, for every \(k\),
The density matrix \(\sigma _\ell \) is independent of \(k\). The tensor factors are written in the reverse order from HJPW Property 1. This is the full-support specialization of that state decomposition. No action of a preserving operation on the summands is asserted in this theorem; the next theorem supplies it under the same full-support hypothesis.
Write the \(\ell \)-th block as \(\sigma _\ell \otimes X_{\ell ,k}\). Positivity of \(\rho _k\), compression to this block, and the partial trace over the first factor show that \(X_{\ell ,k}\) is positive semidefinite. Put \(q_{\ell |k}=\operatorname{tr}(X_{\ell ,k})\). The trace-one identities for \(\rho _k\) and \(\sigma _\ell \) give \(q_{\ell |k}\geq 0\) and \(\sum _\ell q_{\ell |k}=1\). For positive weight, divide \(X_{\ell ,k}\) by its trace. At zero weight, positivity forces \(X_{\ell ,k}=0\), so the maximally mixed density matrix may be chosen. In both cases \(X_{\ell ,k}=q_{\ell |k}\tau _{\ell |k}\).
Use the decomposition of Theorem 13.6.66. For every trace-preserving completely positive operation \(F\) satisfying \(F(\rho _k)=\rho _k\) for all \(k\), and every summand \(\ell \), there is a trace-preserving completely positive operation \(F_\ell \) on \(M_{m_\ell }(\mathbb {C})\) such that \(F_\ell (\sigma _\ell )=\sigma _\ell \). If \(\iota _\ell \) denotes inclusion of the \(\ell \)-th diagonal summand and
then, for all \(A\in M_{m_\ell }(\mathbb {C})\) and \(B\in M_{d_\ell }(\mathbb {C})\),
This is the full-support specialization of HJPW Property \(2'\) (Appendix A, lines 808–816 and 860–882), with the tensor factors reversed. It does not include the joint-support reduction or the Stinespring form of Property 2.
For a preserving operation \(F\), every Kraus operator commutes with the common invariant algebra. In the common direct-sum coordinates it therefore has the form \(\bigoplus _\ell C_{i,\ell }\otimes \mathbf1\). Compressing \(\sum _iC_{i,\ell }^\dagger C_{i,\ell }=\mathbf1\) to a summand gives trace preservation of \(F_\ell \). The displayed action follows by multiplication. Positive definiteness of the common average implies that every summand has positive weight for some \(\rho _k\). Compressing \(F(\rho _k)=\rho _k\) and tracing over the second factor then gives \(F_\ell (\sigma _\ell )=\sigma _\ell \).
Let \(\rho _1,\ldots ,\rho _K\) be density operators and let \(Q\) be the support projection of their average. There are an integer \(r\) and an isometry \(V:\mathbb C^r\to \mathcal H\) with \(VV^\dagger =Q\) such that the compressed states \(\widehat\rho _k=V^\dagger \rho _kV\) are density operators, their average is positive definite, and
Every trace-preserving completely positive operation preserving all \(\rho _k\) compresses along the same isometry to a trace-preserving completely positive operation \(\widehat F\) preserving all \(\widehat\rho _k\). Its ambient and compressed actions intertwine:
This is the reduction to the minimum joint supporting subspace in HJPW, Appendix A, lines 761–763.
The kernel of the average is the intersection of the kernels of the positive summands. Hence \(Q\rho _kQ=\rho _k\) for every \(k\). Choose an isometry onto the range of \(Q\). Compression then preserves positivity and trace, reconstructs each state, and makes the compressed average positive definite. Invariance of the average implies that every Kraus operator preserves its support. Compressing the Kraus operators therefore preserves both trace and every compressed state. Expanding the two Kraus sums and using support invariance gives the displayed intertwining identity.
Let \(\rho _1,\ldots ,\rho _K\) be density operators, without a faithfulness assumption on their average. On their minimum joint supporting subspace there is one direct-sum tensor decomposition in which
where the weights form probability distributions and all displayed factors are density operators. Every trace-preserving completely positive operation preserving the compressed family acts on each diagonal summand as
In particular, every operation preserving the original family restricts to such an operation, and its action transports back through the support isometry by the intertwining identity above. The support isometry reconstructs every original \(\rho _k\) from \(\widehat\rho _k\).
This is HJPW Properties 1 and \(2'\) after the joint-support reduction in Appendix A, lines 761–816 and 853–882. The tensor factors are written in the reverse order from HJPW.
Apply Theorem 13.6.68. The common average is positive definite in the resulting support coordinates, so Theorem 13.6.67 applies there. Apply the full-support block-action theorem to every preserving operation of the compressed family. Operations preserving the original family restrict along the same support isometry, and the intertwining identity transports their action back to the ambient space. Use the reconstruction identity for the family members.
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and let \(\mu _s\) be the finite nonempty family of normalized conditional states obtained from the active separating effects. On the minimum joint supporting subspace of this family there are tensor-product direct-sum coordinates such that
where the weights form probability distributions and all factors are density operators. The support isometry reconstructs every \(\mu _s\).
The channel
restricts to the same support and intertwines with its ambient action. In these coordinates its action on each diagonal summand is
The tensor factors are written in the reverse order from HJPW.
This is the application of HJPW Theorem 6, lines 493–505, to Properties 1 and \(2'\) from Appendix A, lines 761–816 and 853–882.
Theorem 13.6.52 supplies a finite nonempty density family together with a trace-preserving completely positive operation fixing every member. Apply Theorem 13.6.69 to this family. Retain the support reconstruction, the normalized block equations, and the restriction and block-action clauses for the preserving operation.
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and let \(V\) be the support isometry and \((e,U,\sigma _j)\) the direct-sum coordinates of Theorem 13.6.70. Put \(W=\mathbf1_A\otimes V\). Then
In the same support coordinates there are positive, not necessarily normalized operators \(\omega _j\) such that
The tensor factors are written in the reverse order from HJPW. The block equation is on the minimum joint supporting subspace; it does not assert an ambient direct-sum equivalence on subsystem \(B\).
Scope restriction (HJPW Theorem 6, equation (14)): the displayed direct sum is restricted to the minimum joint supporting subspace, whereas equation (14), lines 499–502, decomposes the ambient subsystem \(B\). This restriction is documented in the TNLean paper-gap note [ con26q ] .
Rescale the normalized equations for the active separating effects to their unnormalized conditional slices. A positive slice associated with an inactive effect has trace zero and therefore vanishes. The finite separating family then reconstructs \(\rho _{AB}\) through \(\mathbf1_A\otimes V\). In the transformed support coordinates, take \(\omega _j\) to be the partial trace over the common factor of the \(j\)-th principal block. Positivity follows from positivity of principal blocks and of partial trace. Applying separation once more gives the displayed direct-sum equation, and taking traces gives the normalization.
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity. Then the ambient middle subsystem has a direct-sum tensor decomposition
In unitary coordinates adapted to this decomposition there are density matrices \(\sigma _j\) and positive, not necessarily normalized operators \(\omega _j\) such that
The tensor factors appear here in the order \(b_j^R\otimes b_j^L\), the reverse of HJPW’s \(b_j^L\otimes b_j^R\) order. Directions complementary to the minimum joint support form one-dimensional tensor sectors with zero \(\omega _j\).
This is [ HJPW04 , Theorem 6, equations (13)–(14) ] , lines 493–502.
Extend the support isometry, after composing it with the support block unitary, to a unitary on the ambient middle subsystem. Split the orthogonal complement into one-dimensional tensor sectors. Give each such sector its unique density matrix and the zero unnormalized conditional factor. The supported sectors retain the factors from Theorem 13.6.71. Lifting the unitary extension through subsystem \(A\) and combining the complementary zero block with the supported direct sum gives the displayed ambient equation.
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, with the ambient decomposition
from Theorem 13.6.72. There are a finite-dimensional ancilla, a fixed pure ancilla vector, and one unitary \(U_{BCE}\) whose block-coordinate form is
On every supported sector, \(U_j\) dilates a trace-preserving completely positive map from \(b_j^R\) to \(b_j^R C\) that sends \(\sigma _j\) to a state whose \(b_j^R\) marginal is \(\sigma _j\). The resulting state is positive and has trace one. The displayed block form uses the order \(U_j\otimes \mathbf1_{b_j^L}\); in HJPW’s order it is \(\mathbf1_{b_j^L}\otimes U_j\).
The recovery operation determined by the chosen unitary agrees with the Petz recovery operation on every operator supported by the minimum joint supporting subspace, and
No equality of the two operations is asserted on the complementary ambient sectors.
This is [ HJPW04 , Theorem 6, equation (15) ] , lines 547–560, using Appendix A, Theorem 10, Property 2, lines 791–800, the equivalence with Property \(2'\) in lines 808–823, and the operation-level proof in lines 853–882.
First obtain the ambient decomposition from Theorem 13.6.72. Choose rectangular Kraus operators for the Petz recovery operation and slice them along subsystem \(C\). On the minimum joint supporting subspace, Property \(2'\) makes every slice block diagonal with the identity on \(b_j^L\), while the induced operation on \(b_j^R\) fixes \(\sigma _j\). Extend each sector Stinespring isometry from the same fixed pure ancilla vector to a unitary. Their direct sum gives the block-coordinate unitary; conjugating by the ambient change of basis gives the physical unitary. Equality of the two Stinespring isometries on the support proves agreement with the Petz operation there. Combining this agreement with the supported reconstruction of \(\rho _{AB}\) gives the displayed recovery identity.
Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and let \(U_B\) and the factors \(\omega _j\) be those of Theorems 13.6.72 and 13.6.73. For each supported sector, put
This is a density matrix on \(\mathcal H_{b_j^R}\otimes \mathcal H_C\), where \(\widehat{\mathcal R}_j\) is the sector recovery operation. Read each middle-system summand in the HJPW order \(\mathcal H_{b_j^L}\otimes \mathcal H_{b_j^R}\), and set \(W=\mathbf1_A\otimes (U_B\otimes \mathbf1_C)\). Then
Thus every complementary block is zero, while the supported \(j\)th block, in the factor order \((A b_j^L)\otimes (b_j^R C)\), is \(\omega _j\otimes \widehat\rho _j\). The factors \(\omega _j\) remain unnormalized; no probabilities or normalized left factors are asserted.
This is the final substitution in [ HJPW04 , Theorem 6, equations (11), (14), and (15) ] , lines 562–570.
Equations (11) and (14) give
The physical unitary in equation (15) is the conjugate of its block unitary by \(U_B\otimes \mathbf1_C\). Hence conjugation by \(W\) changes the first identity into the sectorwise action of the block recovery operation on the second identity. In HJPW factor order the supported input block is \(\omega _j\otimes \sigma _j\). On the complement the input block is zero. On a supported sector,
Taking the direct sum proves (177) with the stated orientation \(W^\dagger \rho _{ABC}W\).
Let \(\rho _{ABC}\) be a tripartite density matrix satisfying
Then there are a decomposition
a unitary \(V_B\), probabilities \(p_j\), and density matrices \(\rho _{A B_j^L}\) and \(\rho _{B_j^R C}\) such that, for \(V=\mathbf1_A\otimes V_B\otimes \mathbf1_C\),
By Theorem 13.6.74, put \(W=\mathbf1_A\otimes (U_B\otimes \mathbf1_C)\). Then
where \(\omega _j\geq 0\), each supported sector output state \(\widehat\rho _j\) is positive with trace one, and \(\sum _j\operatorname{tr}\omega _j=1\). Reindex every middle-system fibre from the recovery order \(B_j^R\otimes B_j^L\) to the Hayashi order \(B_j^L\otimes B_j^R\), and choose the Hayashi middle unitary \(V_B=U_B^\dagger \). This gives the required orientation \(V\rho _{ABC}V^\dagger =W^\dagger \rho _{ABC}W\).
For every ambient sector set
Positivity of \(\omega _j\) gives \(p_j\geq 0\), and the total trace identity gives \(\sum _jp_j=1\). Total positive-semidefinite normalization gives a density matrix \(\overline\omega _j\) on \(A B_j^L\) and a density matrix \(\overline\rho _j\) on \(B_j^R C\) such that
On every supported sector, \(\overline\rho _j=\widehat\rho _j\), and \(\overline\omega _j=\omega _j/p_j\) whenever \(p_j{\gt}0\). If a supported sector has \(p_j=0\), its left factor may be any density matrix because \(0\, \overline\omega _j\otimes \widehat\rho _j=0\). On a complementary sector, where both raw factors vanish, either normalized factor may be chosen. These choices are made only at this final boundary, and every weighted summand is unchanged. The displayed direct sum is therefore a quantum Markov decomposition in the sense of Definition 13.6.55.
A tripartite density matrix \(\rho _{ABC}\) that admits a quantum Markov decomposition on the middle subsystem \(B\) satisfies
Write the state in the adapted basis as the block-diagonal direct sum \(\bigoplus _j p_j\, \rho _{A B_j^L}\otimes \rho _{B_j^R C}\). The von Neumann entropy of a weighted orthogonal direct sum is \(S(\bigoplus _j p_j\, \omega _j) =-\sum _j p_j\log p_j+\sum _j p_jS(\omega _j)\) by Theorems 13.4.84 and 13.4.83, and the entropy of a tensor product is additive, \(S(\omega \otimes \tau )=S(\omega )+S(\tau )\), by Theorem 13.4.82. Tracing out one tensor factor within each block, the three reduced states factor as block-diagonal direct sums,
with \(\rho _{ABC}\cong \bigoplus _j p_j\, \rho _{AB_j^L}\otimes \rho _{B_j^R C}\) itself. Writing \(H=-\sum _j p_j\log p_j\), the four entropies expand as
Both sides of the claimed identity equal
so \(S(\rho _{ABC})+S(\rho _B)=S(\rho _{AB})+S(\rho _{BC})\). The basis change on \(B\) and the direct-sum reindexing leave every entropy term unchanged.
For a tripartite density matrix \(\rho _{ABC}\),
holds if and only if \(\rho _{ABC}\) admits a quantum Markov decomposition on the middle subsystem \(B\).
13.7 Mutual information
The quantum mutual information of a bipartite state \(\rho _{AB}\) is
For any bipartite density matrix \(\rho _{AB}\), \(I(A{:}B) \ge 0\).
Apply strong subadditivity (Theorem 13.6.4) with trivial \(B\) (one-dimensional middle system). The SSA inequality \(S(\rho _{ABC}) + S(\rho _B) \le S(\rho _{AB}) + S(\rho _{BC})\) reduces to subadditivity \(S(\rho _{AC}) \le S(\rho _A) + S(\rho _C)\), which gives \(I(A{:}C) \ge 0\).
13.8 Entropy formulations
This section states the entropy formulations used later in the development: von Neumann entropy, strong subadditivity, quantum Markov decomposition, and mutual information. These statements are cited from the standard entropy literature and supply the entropy-theoretic input for the later MPDO arguments.
This formulation has the same value \(S(\rho ) = -\sum _i \lambda _i \log \lambda _i\) as in Definition 13.4.1.
For any tripartite density matrix \(\rho _{ABC}\) on \(A \otimes B \otimes C\),
This formulation introduces no new axiom: it is the same strong-subadditivity statement as Theorem 13.6.4, which is proved there from Lieb concavity along the relative-entropy route [ LR73 ] .
This is the entropy formulation of the Hayashi Markov decomposition from Definition 13.6.55.
For any tripartite density matrix \(\rho _{ABC}\), equality in strong subadditivity holds if and only if \(\rho _{ABC}\) admits a quantum Markov decomposition on the middle subsystem \(B\).
This formulation introduces no new axiom: it is the same equality criterion as Theorem 13.6.77.
This formulation has the same value \(I(A{:}B) = S(\rho _A) + S(\rho _B) - S(\rho _{AB})\) as in Definition 13.7.1.
For a tripartite density matrix \(\rho _{ABC}\) with \(\dim B = 1\), one has \(S(\rho _{ABC}) \le S(\rho _{AB}) + S(\rho _{BC})\).
13.9 Mutual information: monotonicity and area-law bound
The two inequalities below are the downstream MPDO-facing consequences of the entropy inequalities in this chapter. The monotonicity inequality is the strong-subadditivity content underlying the MPDO mutual-information monotonicity \(I_L \le I_{L+1}\) (arXiv:1606.00608, Proposition C.1); the elementary area-law bound is the single-site entropy bound underlying the MPDO area-law bound \(I_L \le 4\log D\) (arXiv:1606.00608, cited from the Wolf area-law bound). Both inequalities follow directly from strong subadditivity (Theorem 13.6.4) and the single-system \(S(\rho ) \le \log D\) bound; neither introduces a new axiom.
For any PSD Hermitian matrix \(\rho \) with \(\operatorname{tr}(\rho ) = 1\) on an arbitrary finite index set, \(S(\rho ) \ge 0\). This is the analog of Theorem 13.4.4 for arbitrary finite index sets: it does not require the index set to be \(\{ 0,\ldots ,D{-}1\} \), so it applies to bipartite density matrices on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B}\).
The eigenvalues of a PSD matrix are non-negative, and eigenvalues of a trace-\(1\) Hermitian matrix with non-negative eigenvalues are bounded above by \(1\) (each single eigenvalue is at most the total sum). Apply \(x\log x \le 0\) on \([0, 1]\) pointwise and sum.
For any tripartite density matrix \(\rho _{ABC}\) on \(A \otimes B \otimes C\),
where the left-hand side is the bipartite mutual information of the reduced state \(\rho _{AB} = \operatorname{tr}_C(\rho _{ABC})\) and the right-hand side is evaluated by expanding \(I(A{:}BC) = S(\rho _A) + S(\rho _{BC}) - S(\rho _{ABC})\).
Expanding both sides in entropy form and cancelling the common \(S(\rho _A)\) term, the inequality reduces to strong subadditivity \(S(\rho _{ABC}) + S(\rho _B) \le S(\rho _{AB}) + S(\rho _{BC})\). The bipartite \(B\)-reduced state of \(\rho _{AB}\) matches the tripartite \(B\)-reduced state \(\operatorname{tr}_{AC}(\rho _{ABC})\), ensuring the \(S(\rho _B)\) term is the one supplied by Theorem 13.8.2.
For a bipartite density matrix \(\rho _{AB}\) on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B}\) with \(d_A, d_B \ge 1\) whose single-system reduced states are obtained by partial trace, \(I(A{:}B) \le \log d_A + \log d_B\).
13.10 Data processing under local channels
Let \(\rho _{AB}\) be a bipartite density operator and let \(\Phi _A\) be a trace-preserving completely positive map whose input and output matrix algebras may have different dimensions. Then \(I(A':B)_{(\Phi _A\otimes \operatorname{id}_B)(\rho )}\leq I(A:B)_\rho \).
Choose a rectangular Stinespring isometry for \(\Phi _A\). Conjugating by \(W=V_A\otimes \operatorname{id}_B\) gives \(\omega _{A'EB}=W\rho _{AB}W^\dagger \). Since \(W=V_A\otimes \operatorname{id}_B\) and \(V_A^\dagger V_A=\operatorname{id}_A\), its marginals satisfy
Entropy is preserved by the isometry \(W\), and also by \(V_A\) on the first marginal. Therefore
Hence \(I(A'E:B)_\omega =I(A:B)_\rho \). The defining Stinespring identity is \(\operatorname{tr}_E(\omega _{A'EB})=(\Phi _A\otimes \operatorname{id}_B)(\rho _{AB})\). Strong subadditivity in the form \(I(R:A')\leq I(R:A'E)\) proves the stated inequality with \(R=B\).
Let \(\rho _{AB}\) be a bipartite density operator and let \(\Psi _B\) be a trace-preserving completely positive map whose input and output matrix algebras may have different dimensions. Then \(I(A:B')_{(\operatorname{id}_A\otimes \Psi _B)(\rho )}\leq I(A:B)_\rho \).
Exchange the two tensor factors and apply Theorem 13.10.1.
Let \(\rho _{AB}\) be a bipartite density operator and let \(\Phi _A\) and \(\Psi _B\) be trace-preserving completely positive maps whose input and output matrix algebras may have different dimensions. Then \(I(A':B')_{(\Phi _A\otimes \Psi _B)(\rho )}\leq I(A:B)_\rho \).
Apply the two one-sided inequalities successively.
13.11 Classical information and operator-Schmidt bounds
Let \(X\) and \(Y\) be finite sets. A matrix \(P=(P_{x,y})_{x\in X,y\in Y}\) is a joint probability distribution if \(P_{x,y}\geq 0\) for every \(x\in X\) and \(y\in Y\), and
Its row and column marginals are respectively
For \(t{\gt}0\), set \(h(t)=-t\log t\), and set \(h(0)=0\). The entropy of a probability distribution \(a=(a_z)_{z\in Z}\) on a finite set is \(H(a)=\sum _{z\in Z}h(a_z)\). For a joint probability distribution \(P\) with row and column marginals \(p\) and \(q\), put
Let \(p=(p_z)_{z\in Z}\) be a probability distribution on a finite set \(Z\). Define \(h(t)=-t\log t\) for \(t{\gt}0\) and \(h(0)=0\). Then
Let \(k=\lvert \{ z\in Z:p_z\neq 0\} \rvert \). Jensen’s inequality for the concave function \(h\), with uniform weights on the support, gives
Multiplication by \(k\) proves the result, since values outside the support contribute \(h(0)=0\).
Let \(X\) and \(Y\) be finite sets and let \(P=(P_{x,y})_{x\in X,y\in Y}\) be a joint probability distribution. With natural logarithms, \(I(X:Y)_P\leq \log \operatorname{rank}_{\mathbb {R}}P\). Equivalently, with logarithms to base two, \(2^{I(X:Y)_P}\leq \operatorname{rank}_{\mathbb {R}}P\).
This is Theorem 4.1 of [ RV17 ] . If the number of nonzero rows exceeds the ordinary real rank, choose a nontrivial real linear relation among those rows. Nonnegativity forces the relation to have coefficients of both signs. Rescaling each row by \(1+\varepsilon \beta _x\) gives two endpoint distributions. For every \(y\in Y\), the row relation gives
Thus each endpoint preserves every column marginal, removes at least one row, and has no larger rank. Write \(h(t)=-t\log t\) for \(t{\gt}0\), with \(h(0)=0\), and put \(m=\min _x\beta _x{\lt}0{\lt}M=\max _x\beta _x\). Then
If \(S(\beta )\geq 0\), choose \(\varepsilon =-m^{-1}\geq 0\); otherwise choose \(\varepsilon =-M^{-1}\leq 0\). In either case \(\varepsilon S(\beta )\geq 0\), so the selected endpoint has mutual information no smaller than the original distribution. Repeating the construction leaves at most \(\operatorname{rank}_{\mathbb {R}}P\) nonzero rows. Finally, \(I(X:Y)\leq H(X)\), and the entropy of a distribution supported on \(k\) points is at most \(\log k\).
13.11.1 Operator-Schmidt rank and marginal-support compression
Every nonzero finite-dimensional bipartite operator \(X\) satisfies
This includes zero-dimensional factors, where the nonzero hypothesis is impossible.
If \(\operatorname {OSR}(X)=0\), a shortest product decomposition of \(X\) is an empty sum. Hence \(X=0\).
Write \(\rho _{ij}\in M_{d_B}(\mathbb {C})\) for the blocks determined by a basis of the first factor, and define \(\mathcal R_\rho (X)=\sum _{i,j}X_{ij}\rho _{ij}\). Then
The minimality theorem at the preceding definition proves that the range dimension \(\dim \operatorname{range}\mathcal R_\rho \) is the least admissible product-decomposition length: it is itself an admissible length and no shorter length suffices. Since the operator-Schmidt rank is defined as the least such length, uniqueness of the least integer identifies the two, giving the first equality \(\operatorname {OSR}(\rho )=\dim \operatorname{range}\mathcal R_\rho \). For the second equality, \(\mathcal R_\rho \) is the linear combination map sending a coefficient matrix \(X\) to \(\sum _{i,j}X_{ij}\rho _{ij}\), so its range is exactly the span of the operator blocks \(\rho _{ij}\).
Let \(V_A:\mathcal H_A\to \mathcal K_A\) and \(V_B:\mathcal H_B\to \mathcal K_B\) be isometries between finite-dimensional complex spaces. For every operator \(X\) on \(\mathcal H_A\otimes \mathcal H_B\),
This remains valid when one or more of the spaces have dimension zero.
More generally, multiplying on the left and right by product matrices cannot increase operator-Schmidt rank: apply the four local matrices to the two factors in each term of a shortest product decomposition. Applying this inequality first to \(V_A,V_B\) and then to their adjoints gives the two inequalities. The identities \(V_A^\dagger V_A=1\) and \(V_B^\dagger V_B=1\) identify the twice-transformed operator with \(X\).
Let \(\rho _{AB}\succeq 0\), with marginals \(\rho _A\) and \(\rho _B\). Then
Equivalently, the support of \(\rho _{AB}\) is contained in \(\operatorname{supp}(\rho _A)\otimes \operatorname{supp}(\rho _B)\). This remains valid when either factor has dimension zero.
Let \(P_A\) and \(P_B\) be the support projections of the marginals. The two right-absorption identities give \(\rho _{AB}(P_A\otimes P_B)=\rho _{AB}\). The support projection of \(\rho _A\otimes \rho _B\) is \(P_A\otimes P_B\). Therefore every vector annihilated by \(\rho _A\otimes \rho _B\) is annihilated by \(\rho _{AB}\).
Let \(\rho \geq 0\) be an operator on \(H_A\otimes H_B\). Let \(V_A:\widehat H_A\to H_A\) and \(V_B:\widehat H_B\to H_B\) be isometries whose range projectors are the support projectors \(P_A\) and \(P_B\) of the two marginals. Set \(W=V_A\otimes V_B\) and \(\rho _c=W^\dagger \rho W\). Then
No marginal is required to be faithful, and zero-dimensional coordinate spaces are permitted.
The four marginal-support absorption identities give
Since \(WW^\dagger =P_A\otimes P_B\), associativity yields
The local-isometry theorem applied to \(\rho _c\) gives \(\operatorname {OSR}(W\rho _cW^\dagger )=\operatorname {OSR}(\rho _c)\). Substitution of the reconstruction identity proves the rank equality.
Let \(A:H_A\to K_A\) and \(B:H_B\to K_B\) be linear maps between finite-dimensional complex spaces, and let \(X\) be an operator on \(H_A\otimes H_B\). If \(B^\dagger B=1\), then
If \(A^\dagger A=1\), then, symmetrically,
The identities include zero-dimensional spaces.
Expand the matrix entries in orthonormal bases. In the first identity, the sum over the traced output basis gives the matrix entries of \(B^\dagger B=1\), and hence collapses the two input indices on \(H_B\). The second identity follows by the same argument with the factors interchanged.
In the setting of Theorem 13.11.1.5, the two marginals of \(\rho _c\) satisfy
Both compressed marginals are positive definite, including when one of their spaces has dimension zero.
Insert the reconstruction \(\rho =W\rho _cW^\dagger \) into the two covariance identities of Lemma 13.11.1.6, and multiply by the corresponding adjoint isometries. Each displayed marginal is the compression of a positive semidefinite matrix to its support, so its positive definiteness follows from Theorem 10.4.1.5.
Let \(\rho \geq 0\) be a bipartite complex matrix whose first marginal is faithful. Choose an eigenbasis in which \(\operatorname{tr}_B\rho =\operatorname{diag}(p_1,\ldots ,p_{d_A})\) with every \(p_i{\gt}0\), and define the linear map \(\Phi _\rho \) on matrix units by
Write \(\sigma =\operatorname{diag}(p_1,\ldots ,p_{d_A})\) and \(\tau =\operatorname{tr}_A\rho \). Then \(\Phi _\rho \) is completely positive and trace preserving,
and application of \(\Phi _\rho \) to the first half of \(\sum _{i,j}\sqrt{p_ip_j}\, E_{ij}\otimes E_{ij}\), followed by restoring the order of the two factors, reconstructs \(\rho \).
Strict positivity of the \(p_i\) makes the entrywise input scaling invertible, so it does not change the range. Positivity of \(\rho \) gives a Kraus representation of \(\Phi _\rho \), while the first marginal equation gives trace preservation. Summing the diagonal blocks gives \(\Phi _\rho (\sigma )=\tau \). Substitution on matrix units proves the reconstruction identity. This is the finite-dimensional Choi representation [ Cho75 ] with the input marginal absorbed into the canonical purification.
13.12 Support compression for entropy functionals
Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) be an isometry, so that \(V^\dagger V=\mathbb {1}_k\). Define \(\iota _V:M_{k}(\mathbb {C})\to M_{D}(\mathbb {C})\) by
The map \(\iota _V\) is complex-linear, multiplicative, and \(*\)-preserving. In general it is not unital: \(\iota _V(\mathbb {1}_k)=VV^\dagger \) is the projection onto the range of \(V\).
Let \(A\in M_{k}(\mathbb {C})\) be Hermitian, let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy \(V^\dagger V=\mathbb {1}_k\), and let \(f:\mathbb {R}\to \mathbb {R}\) satisfy \(f(0)=0\). Then
where both sides use the continuous functional calculus on the respective finite spectra.
Apply functoriality of the continuous functional calculus to \(\iota _V\). The condition \(f(0)=0\) removes the contribution from the orthogonal complement of the range of \(V\), where \(\iota _V(A)\) vanishes.
Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy \(V^\dagger V=\mathbb {1}_k\). For every \(A\in M_{k}(\mathbb {C})\),
Cyclicity of the trace gives \(\operatorname{tr}(VAV^\dagger )=\operatorname{tr}(V^\dagger VA)=\operatorname{tr}(A)\).
Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy \(V^\dagger V=\mathbb {1}_k\). For every \(A\in M_{k}(\mathbb {C})\),
Associativity and \(V^\dagger V=\mathbb {1}_k\) reduce the left-hand side to \((V^\dagger V)A(V^\dagger V)=A\).
Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite and suppose that \(\ker \omega \subseteq \ker \rho \). If \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfies \(VV^\dagger =P_{\operatorname{supp}(\omega )}\), then
No condition on \(V^\dagger V\) is needed for this identity.
The kernel inclusion and positivity give \(\rho P_{\operatorname{supp}(\omega )}=\rho \); taking adjoints also gives \(P_{\operatorname{supp}(\omega )}\rho =\rho \). Substituting \(VV^\dagger =P_{\operatorname{supp}(\omega )}\) on both sides of \(\rho \) proves (187).
Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite with \(\ker \omega \subseteq \ker \rho \). Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy
Then
The logarithm is totalized by \(\log 0=0\) on the kernel.
If \(A\in M_{n}(\mathbb {C})\) is Hermitian and \(f:\mathbb {R}\to \mathbb {R}\), then
In particular,
If \(A\) is positive semidefinite, then for every \(r\in \mathbb {R}\),
These identities include singular matrices and matrices whose index set is empty. They are project-derived rather than statements of CPSV16.
For a Hermitian matrix, transpose is entrywise complex conjugation. Apply covariance of the continuous functional calculus under this real star-algebra automorphism. Continuity of \(f\) is needed only on the finite spectrum of \(A\), where it is automatic. The logarithm, real-power, and square-root identities follow by specialization. For the support projection, specialize to \(f(x)=1\) for \(x\neq 0\) and \(f(0)=0\).
If \(\rho \succeq 0\), then
On every positive eigenvalue this is the scalar identity \(\log \sqrt{x}=\frac12\log x\). At \(x=0\), both sides vanish under the convention \(\log 0=0\).
If \(A\succ 0\), then, for every \(s\in \mathbb {R}\),
Every eigenvalue of \(A\) is positive, so the assertion follows from \(\log (x^s)=s\log x\) on \((0,\infty )\) and the continuous functional calculus.
Let \(A\in M_{n}(\mathbb {C})\) be positive semidefinite and let \(v\in \mathbb {C}^n\) satisfy \(\langle v,v\rangle =1\) and \(P_{\operatorname{supp}(A)}v=v\). Then
Here \(\log A\) is defined by the continuous functional calculus with \(\log 0=0\). Compare the finite-dimensional vector-state Jensen argument in [ Bha97 , Chapter V ] .
Diagonalize \(A=U\operatorname{diag}(\mu _i)U^\dagger \), set \(w=U^\dagger v\), and put \(p_i=|w_i|^2\). The support condition implies \(\mu _i=0\Longrightarrow w_i=0\), so every spectral weight at zero vanishes. Thus \(p_i\geq 0\) for every \(i\in S=\{ i:\mu _i{\gt}0\} \) and \(\sum _{i\in S}p_i=1\). The spectral formulas give
Scalar concavity of the logarithm on \((0,\infty )\) now gives (196). No concavity assertion at zero and no positive-definiteness of \(A\) are used.
For square matrices \(\rho \) and \(\omega \) of the same size, define
The negative power is defined by functional calculus and vanishes on the kernel of \(\omega \). On trace-one positive semidefinite matrices satisfying \(\ker \omega \subseteq \ker \rho \), its logarithm is the order-two sandwiched Rényi divergence [ MLDS\(^{+}\)13 , Definition 2 ] .
If \(\rho \succeq 0\) and \(\omega \succeq 0\), then \(Q_2(\rho ,\omega )\geq 0\).
The sandwiched matrix \(\omega ^{-1/4}\rho \, \omega ^{-1/4}\) is positive semidefinite. The trace of its square is therefore nonnegative.
If \(\omega \succ 0\), then, for every square matrix \(\rho \) of the same size,
Combine the two adjacent inverse quarter-powers by \(\omega ^{-1/4}\omega ^{-1/4}=\omega ^{-1/2}\) and rotate the factors under the trace. No Hermiticity assumption on \(\rho \) is needed.
Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) satisfy \(\rho \succeq 0\), \(\omega \succ 0\), and \(\operatorname{tr}\rho =1\). Then
No normalization of \(\omega \) is required.
Set \(R=\sqrt\rho \), \(T=\omega ^{-1/2}\), \(\Delta =R^T\otimes T\), and \(v=\operatorname {vec}(R)\). The vectorization is column-stacking, with factor order
Hence \(\Delta v=\operatorname {vec}(T\rho )\), and cyclicity of the trace identifies its first moment with
The support projection of \(\Delta \) fixes \(v\):
Indeed, \(T\) is positive definite, so \(P_{\operatorname{supp}(T)}=1\), while the support projection of \(R^T\) fixes \(R^T\) and hence \(R P_{\operatorname{supp}(R^T)}^T=R\). The column-stacking identity (200) proves the displayed equality. This support condition is essential when \(\rho \) is singular: it removes the zero spectral weights before Jensen’s inequality is applied.
Theorem 13.4.70, 13.12.8, 13.12.9, and 13.12.7 give
Theorem 13.12.10 therefore yields \(\frac12D(\rho \Vert \omega )\leq \log m\). Since \(\| R\| _{\mathrm{HS}}^2=\operatorname{tr}\rho =1\), Hilbert–Schmidt Cauchy–Schwarz gives
Positivity of \(m\) now gives \(D(\rho \Vert \omega )\leq \log (m^2)\leq \log Q_2(\rho ,\omega )\).
The auxiliary divergence and its vector-state Jensen argument are adapted from [ MLDS\(^{+}\)13 , arXiv:1306.3142v4, Definition 5 and Lemma 19, used in the proof of Theorem 7 ] . Only the direct order-two endpoint is used here; the final Cauchy–Schwarz estimate is not an application of Beigi’s interpolation theorem.
Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite with \(\ker \omega \subseteq \ker \rho \). Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy (189). Then
The function \(x\mapsto x^{-1/4}\) is taken to be zero at \(x=0\), so Theorem 13.12.2 applies. Reconstruct \(\rho \) and \(\omega \), combine the three expanded factors by multiplicativity of \(\iota _V\), and use trace preservation.
Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite matrices of trace one with \(\ker \omega \subseteq \ker \rho \). Assume that for every dimension and every pair of trace-one matrices \(A\succeq 0\) and \(B\succ 0\),
Then
Thus this theorem reduces the singular-reference case to the faithful comparison. Theorem 13.12.14 supplies the hypothesis in (206).
Choose an isometric inclusion of the support of \(\omega \). Its compression \(\omega _V=V^\dagger \omega V\) is positive definite by Theorem 10.4.1.5, while \(\rho _V=V^\dagger \rho V\) remains positive semidefinite. Reconstruction and (185) show that both compressed matrices have trace one. Apply (206) to \((\rho _V,\omega _V)\), then use (190) and (205).
Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite matrices of trace one. If \(\ker \omega \subseteq \ker \rho \), then
13.12.1 Whitened Choi estimates and sandwiched Rényi bounds
Let \(\rho _{AB}\geq 0\), and choose an eigenbasis of its faithful first marginal \(\sigma =\operatorname{diag}(p_1,\ldots ,p_{d_A}){\gt}0\). Suppose also that the second marginal \(\tau \) is positive definite. For the supported-marginal channel \(\Phi _\rho \), set
Then
Here \(\sigma ^T=\sigma \) because the chosen basis diagonalizes \(\sigma \).
For arbitrary congruence matrices \(A\) and \(B\), the matrix-unit convention for the Choi matrix gives
Apply this identity to the inverse-square-root scaling in \(\Phi _\rho \) and to the two fourth-root factors defining \(L\). Entrywise diagonal functional calculus gives \(\sigma ^{1/4}=\operatorname{diag}(p_i^{1/4})\) and \(\operatorname{diag}(p_i^{-1/2})\sigma ^{1/4}=\sigma ^{-1/4}\), proving the first assertion. The two modular congruences are invertible linear maps, so pre- and post-composition by them preserve the range dimension.
Under the hypotheses of Theorem 13.12.1.1, the matrix
is positive semidefinite, and
This is the faithful-marginal-support form of the order-two estimate obtained from [ Bei13 , Theorem 6, Equation (18) ] .
The matrix \(W\) is a congruence of \(\rho _{AB}\geq 0\), so it is positive semidefinite and Hermitian; consequently \(\operatorname{tr}(W^2)=\| W\| _F^2\). The weighted map \(L\) is a Hilbert–Schmidt contraction. Its Choi matrix is the whitened state by the preceding theorem, while invertibility of the modular congruences and the supported-marginal correspondence give
The rectangular Choi range-dimension estimate now proves the claim.
Let \(A\succeq 0\), let \(s\in \mathbb {R}\), and let \(U\) be unitary. Then
Apply Lemma 13.4.22 to the function \(x\mapsto x^s\).
Let \(A\succeq 0\) and \(B\succeq 0\). For every \(s\in \mathbb {R}\),
Apply Theorem 13.4.42 to \(x\mapsto x^s\), using \((xy)^s=x^sy^s\) for \(x,y\geq 0\).
If \(A\succ 0\), then
The continuous functional calculus gives \(\sqrt{\sqrt A}=A^{1/4}\). Positive definiteness makes this matrix invertible, and inversion changes the exponent from \(1/4\) to \(-1/4\).
Let \(\omega \succeq 0\), let \(\rho \) be a square matrix of the same size, and let \(U\) be unitary. Then
Theorem 13.12.1.3 gives \((U\omega U^\dagger )^{-1/4}=U\omega ^{-1/4}U^\dagger \). Substitute this identity into both sandwiching factors and use \(U^\dagger U=1\) and cyclicity of the trace.
Let \(\rho \) be a square matrix and let \(\omega \succeq 0\) have the same finite index set. Let \(e\) be a bijection onto another finite index set. Then
Reindexing commutes with the continuous functional calculus, matrix multiplication, and the trace. Apply these three identities to the two inverse quarter-powers and the two copies of the sandwiched matrix.
Let \(\rho _{AB}\succeq 0\) have positive definite marginals \(\rho _A\) and \(\rho _B\). Then
No trace normalization is required.
Choose a unitary \(U\) that diagonalizes the first marginal and conjugate \(\rho _{AB}\) by \(U^\dagger \otimes 1\). The two marginals become \(\operatorname{diag}(p_i)\) and \(\rho _B\), while positivity and operator-Schmidt rank are unchanged. The transpose in the Choi whitening convention becomes invisible only in this eigenbasis, since \(\operatorname{diag}(p_i)^T=\operatorname{diag}(p_i)\).
Theorems 13.12.1.4 and 13.12.1.5 identify the negative quarter-power of the product marginal with the two whitening factors in Theorem 13.12.1.2. That theorem bounds the squared whitened trace by the ordinary operator-Schmidt rank. Finally, Theorem 13.12.1.6 and the invariance of operator-Schmidt rank under local unitaries return to the original basis.
Let \(\rho _{AB}\succeq 0\) be any finite-dimensional bipartite operator. Then
No trace normalization or positive-dimension hypothesis is required. The statement includes the cases in which either marginal support has dimension zero.
Compress both factors simultaneously to the supports of \(\rho _A\) and \(\rho _B\). The compressed marginals are positive definite by Theorem 13.11.1.7, even when a support coordinate space has dimension zero. The order-two functional is unchanged by Theorem 13.12.15, and the ordinary operator-Schmidt rank is unchanged by Theorem 13.11.1.5. Apply Theorem 13.12.1.8 to the compressed operator.
The sandwiched Rényi comparison also uses the following unnormalized trace term. For trace-one \(\rho \) and \(\alpha {\gt}1\), it is the quantity inside the logarithm of the sandwiched Rényi divergence [ Bei13 , Equation (3) ] [ MLDS\(^{+}\)13 , Definition 2 ] whenever \(\rho ,\omega \geq 0\) and \(\ker \omega \subseteq \ker \rho \). This includes singular \(\omega \), with negative powers interpreted as the generalized inverse on its support. In this regime, finiteness of the divergence requires the support inclusion. The present application uses \(1{\lt}\alpha \leq 2\); the order-one case is the Umegaki endpoint. The statements below concern only this algebraic expression. They are not assertions of [ CPGSV16 , Proposition 4.5 ] , and they do not establish continuity or monotonicity in the order, interpolation between the endpoints, differentiability at order one, a limit to Umegaki relative entropy, a logarithmic divergence, or a comparison with relative entropy.
For square complex matrices \(\rho \) and \(\omega \) and every \(\alpha \in \mathbb {R}\), set
Real powers are defined by continuous functional calculus, with negative powers equal to zero on the zero eigenspace. Thus \(\widetilde Q_\alpha \) is total even when \(\alpha =0\) or \(\omega \) is singular. For \(\alpha {\gt}1\), on positive semidefinite inputs satisfying \(\ker \omega \subseteq \ker \rho \), this convention gives the finite sandwiched Rényi trace term. In this regime, if the support inclusion fails, the totalized value remains finite but is not the divergence trace term.
If \(\rho \geq 0\) and \(\omega {\gt}0\), then, for every \(\alpha \in \mathbb {R}\),
Put \(q=\omega ^{(1-\alpha )/(2\alpha )}\). Continuous functional calculus gives \(q\geq 0\), hence \(q\rho q\geq 0\) and \((q\rho q)^\alpha \geq 0\). The trace of the last matrix is real and nonnegative.
If \(\rho \geq 0\) and \(\omega {\gt}0\), then
At \(\alpha =1\) the two powers of \(\omega \) have exponent zero, while the remaining real power of \(\rho \) has exponent one. Substitution into (212) gives the identity.
If \(\rho \geq 0\) and \(\omega {\gt}0\), then
At \(\alpha =2\) the sandwiching exponent is \(-1/4\). Since \(\omega ^{-1/4}\rho \, \omega ^{-1/4}\) is positive semidefinite, its continuous-functional-calculus square is its ordinary matrix square.
Let \(\rho _{AB}\succeq 0\) be a finite-dimensional bipartite operator with \(\operatorname{tr}\rho _{AB}=1\). Then
No positive-dimension or faithful-marginal hypothesis is required. The operator-Schmidt rank is the ordinary operator-Schmidt rank.
Set \(\omega =\rho _A\otimes \rho _B\). The marginals are positive semidefinite, \(\operatorname{tr}\omega =1\), and Theorem 13.11.1.4 gives \(\ker \omega \subseteq \ker \rho _{AB}\). Reindex the product basis by a single finite coordinate set. Lemma 13.4.33 and Theorem 13.12.1.7 preserve the two functionals, so Theorem 13.12.17 applies. Theorem 13.6.46 identifies its left side with mutual information. Thus, with \(q=Q_2(\rho _{AB},\omega )\),
The second inequality is Theorem 13.12.1.9. Since \(\operatorname{tr}\rho _{AB}=1\), the operator is nonzero, and Theorem 13.11.1.1 gives \(\operatorname {OSR}(\rho _{AB}){\gt}0\). This positive integer is at least one, so its logarithm is nonnegative. If \(q=0\), the totalized convention \(\log 0=0\) gives \(I(A:B)_\rho \leq 0\leq \log \operatorname {OSR}(\rho _{AB})\). If \(q{\gt}0\), monotonicity of the logarithm gives \(\log q\leq \log \operatorname {OSR}(\rho _{AB})\).
13.13 Trivial-factor corollaries
The partial traces and von Neumann entropy are introduced in Chapter 13. This supplement collects elementary trace-preservation identities and the dimension-1 entropy bound that follow directly from those definitions. They are listed as separate results because each proof requires only unfolding definitions and re-indexing finite sums.
For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_A(\rho _{ABC}))\).
Unfolding the definitions,
For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_C(\rho _{ABC}))\).
Unfolding the definitions,
For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_{AC}(\rho _{ABC}))\).
Unfolding the definitions and interchanging the outer sums,
If \(\rho \in M_{1}(\mathbb {C})\) is Hermitian and \(\operatorname{tr}(\rho )=1\), then \(S(\rho )=0\).
The unique eigenvalue equals the trace, hence equals \(1\), and \(-1\cdot \log 1=0\):