21 Quantum Entropy
This chapter develops trace distance, von Neumann entropy, and their basic properties for finite-dimensional quantum systems.
21.1 Trace norm
For a matrix \(A\in M_{D}(\mathbb {C})\), the singular values \(s_0(A),s_1(A),\ldots \) form a finitely supported family: only finitely many are nonzero. The Schatten one-norm of \(A\) is their sum, that is, the sum of the finitely many nonzero singular values:
This is the \(p=1\) case of the Schatten \(p\)-norm [ Wol12 , Chapter 8, Section 8.1 ] . The equivalent closed formula \(\lVert A\rVert _1=\sum _{i=0}^{D-1}s_i(A)\), summing over all \(D\) singular values including trailing zeros, is part of Theorem 21.1.4.
The trace norm of \(A\in M_{D}(\mathbb {C})\) is its Schatten one-norm:
Wolf records the equivalent formula \(\lVert A\rVert _1=\operatorname{tr}\lvert A\rvert \) [ Wol12 , Chapter 8, Section 8.1 ] ; this is Theorem 21.1.5.
The Schatten one-norm and trace norm are the sums over the finite support of the singular-value sequence. Equivalently, the trace norm is the sum over the indices below the rank of the represented linear map.
The singular-value sequence satisfies \(s_i(A)=0\) for all \(i\geq \operatorname{rank}(A)\), so summing over the support and summing over \(\{ 0,\ldots ,\operatorname{rank}(A)-1\} \) give the same value.
For every \(A\in M_{D}(\mathbb {C})\),
Also, \(\lVert A\rVert _{\operatorname{tr}}=0\) if and only if \(A=0\). Equivalently, \(\lVert A\rVert _{\operatorname{tr}}{\gt}0\) if and only if \(A\neq 0\).
The singular values satisfy \(s_i(A)\geq 0\), so \(\lVert A\rVert _{\operatorname{tr}}=0 \iff \forall i,\, s_i(A)=0 \iff A=0\).
For every \(A\in M_{D}(\mathbb {C})\), with \(\lvert A\rvert =\sqrt{A^\dagger A}\) the positive-semidefinite square root of \(A^\dagger A\),
This is the formula \(\lVert A\rVert _1=\operatorname{tr}[\lvert A\rvert ]\) of [ Wol12 , Chapter 8, Section 8.1 ] ; the trace of the positive-semidefinite matrix \(\lvert A\rvert \) is real.
The map \(T^\dagger T\), for \(T\) the linear map on \(\mathbb {C}^D\) represented by \(A\), is represented by \(A^\dagger A\), so the singular values of \(A\) are \(s_i(A)=\sqrt{\lambda _i}\) with \(\lambda _0,\ldots ,\lambda _{D-1}\) the eigenvalues of \(A^\dagger A\). Diagonalizing \(A^\dagger A=U\operatorname{diag}(\lambda _i)U^\dagger \) gives \(\lvert A\rvert =U\operatorname{diag}(\sqrt{\lambda _i})U^\dagger \), hence
where the last step uses that \(\lvert A\rvert \) is positive semidefinite, so \(\operatorname{tr}\lvert A\rvert \geq 0\) is real.
For every \(c\in \mathbb {C}\) and \(A\in M_{D}(\mathbb {C})\), \(\lVert cA\rVert _{\operatorname{tr}} =\lvert c\rvert \, \lVert A\rVert _{\operatorname{tr}}\). This is the homogeneity axiom for matrix norms [ Wol12 , Chapter 8, Section 8.1 ] .
From \((cA)^\dagger (cA)=\lvert c\rvert ^2\, A^\dagger A\) and uniqueness of the positive-semidefinite square root, \(\lvert cA\rvert =\lvert c\rvert \, \lvert A\rvert \). Linearity of the trace gives \(\operatorname{tr}\lvert cA\rvert =\lvert c\rvert \, \operatorname{tr}\lvert A\rvert \), and the claim follows from Theorem 21.1.5.
For all unitaries \(U,V\in M_{D}(\mathbb {C})\) and every \(A\in M_{D}(\mathbb {C})\),
The trace norm is thus unitarily invariant [ Wol12 , Chapter 8, Section 8.1 ] .
From \((UA)^\dagger (UA)=A^\dagger U^\dagger UA=A^\dagger A\) the absolute values agree: \(\lvert UA\rvert =\lvert A\rvert \). For the right factor, \((AV)^\dagger (AV)=V^\dagger (A^\dagger A)V\), and since \(V^\dagger \lvert A\rvert V\) is positive semidefinite with
uniqueness of the positive-semidefinite square root gives \(\lvert AV\rvert =V^\dagger \lvert A\rvert V\). Cyclicity of the trace then yields \(\operatorname{tr}\lvert AV\rvert =\operatorname{tr}(\lvert A\rvert \, VV^\dagger ) =\operatorname{tr}\lvert A\rvert \). For the two-sided case, \(\lVert UAV\rVert _{\operatorname{tr}} =\lVert UA\rVert _{\operatorname{tr}} =\lVert A\rVert _{\operatorname{tr}}\) by applying the right and left cases in turn.
For every \(A\in M_{D}(\mathbb {C})\), the trace norm is the sum of the square roots of the eigenvalues of the positive operator \(A^\dagger A\):
Since \(s_i(A)=\sqrt{\lambda _i(A^\dagger A)}\) by the definition of singular value, the finite singular-value expansion (??) gives
For every \(A\in M_{D}(\mathbb {C})\), with \(\lambda _0,\ldots ,\lambda _{D-1}\) the eigenvalues of the matrix \(A^\dagger A\), \(\lVert A\rVert _{\operatorname{tr}} =\sum _{i=0}^{D-1}\sqrt{\lambda _i}\).
Diagonalizing \(A^\dagger A=V\operatorname{diag}(\lambda _i)V^\dagger \) gives \(\lvert A\rvert =V\operatorname{diag}(\sqrt{\lambda _i})V^\dagger \), so \(\operatorname{tr}\lvert A\rvert =\sum _{i=0}^{D-1}\sqrt{\lambda _i}\), and the claim follows from Theorem 21.1.5.
For every \(A\in M_{D}(\mathbb {C})\),
and the maximum is attained: some unitary \(U\) satisfies \(\operatorname{tr}[A^\dagger U]=\lVert A\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Eq. (8.11) ] .
For the upper bound, let \(v_0,\ldots ,v_{D-1}\) be an orthonormal eigenbasis of \(A^\dagger A\) with eigenvalues \(\lambda _i\), and let \(U\) be unitary. Expanding the trace in this basis and applying the Cauchy–Schwarz inequality,
since \(\lVert Av_i\rVert ^2=\langle v_i,\, A^\dagger A\, v_i\rangle =\lambda _i\) and \(\lVert Uv_i\rVert =1\).
For attainment, the vectors \(w_i=Av_i/\sqrt{\lambda _i}\), taken over the indices with \(\lambda _i\neq 0\), satisfy
so they form an orthonormal family. Extend it to an orthonormal basis \((w_i)_i\) of \(\mathbb {C}^D\) and let \(U\) be the unitary with \(Uv_i=w_i\) for all \(i\). Then
where the indices with \(\lambda _i=0\) contribute \(0\) to both sides because \(\lVert Av_i\rVert ^2=\lambda _i=0\) forces \(Av_i=0\).
For all \(A,B\in M_{D}(\mathbb {C})\), \(\lVert A+B\rVert _{\operatorname{tr}} \leq \lVert A\rVert _{\operatorname{tr}}+\lVert B\rVert _{\operatorname{tr}}\). Together with homogeneity (Theorem 21.1.6) and definiteness (Theorem 21.1.4), this completes the norm axioms of [ Wol12 , Chapter 8, Section 8.1 ] for the trace norm.
Choose by Theorem 21.1.10 a unitary \(U\) with \(\operatorname{tr}[(A+B)^\dagger U]=\lVert A+B\rVert _{\operatorname{tr}}\). Splitting the trace and applying the upper-bound half of the same theorem to \(A\) and to \(B\) separately,
Let \(H\in M_{D}(\mathbb {C})\) be Hermitian with Jordan decomposition \(H=H^+-H^-\) into orthogonal positive parts, \(H^\pm \geq 0\) and \(H^+H^-=0\). Then \(\lVert H\rVert _{\operatorname{tr}}=\operatorname{tr}[H^+]+\operatorname{tr}[H^-]\).
The sum \(H^++H^-\) is positive semidefinite, and by orthogonality its square is
so \(H^++H^-=\sqrt{H^\dagger H}=\lvert H\rvert \). Hence \(\lVert H\rVert _{\operatorname{tr}} =\operatorname{tr}\lvert H\rvert =\operatorname{tr}[H^+]+\operatorname{tr}[H^-]\) by Theorem 21.1.5.
Let \(A\in M_{D}(\mathbb {C})\) be Hermitian with positive part \(A^+\). There is a matrix \(\Pi \) with \(0\leq \Pi \leq \mathbb {1}\), \(\Pi ^2=\Pi \), and \(\Pi A=A^+\), namely the orthogonal projection onto the support space of \(A^+\).
Diagonalize \(A=U\operatorname{diag}(\lambda _1,\ldots ,\lambda _D)\, U^\dagger \) and set
where \(\chi \) is the indicator function of \((0,\infty )\). Each claim is read off eigenvalue-wise: \(0\leq \chi \leq 1\) gives \(0\leq \Pi \leq \mathbb {1}\), \(\chi ^2=\chi \) gives \(\Pi ^2=\Pi \), and \(\chi (\lambda ) \lambda =\max (\lambda ,0)\) gives \(\Pi A=A^+\).
Let \(X,C\in M_{D}(\mathbb {C})\) with \(X\geq 0\). If \(C\geq 0\), then \(\operatorname{tr}[CX]\geq 0\); if \(C\leq \mathbb {1}\), then \(\operatorname{tr}[CX]\leq \operatorname{tr}[X]\).
The first bound is Lemma C.2.1. For the second, \(\operatorname{tr}[X]-\operatorname{tr}[CX]=\operatorname{tr}[(\mathbb {1}-C)X]\geq 0\) by the same lemma, since \(\mathbb {1}-C\geq 0\).
If \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) is a positive linear map and \(H\in M_{D}(\mathbb {C})\) is Hermitian, then \(T(H)\) is Hermitian.
Write \(H=H^+-H^-\) with \(H^\pm \geq 0\). Then \(T(H)=T(H^+)-T(H^-)\) is a difference of positive semidefinite matrices, hence Hermitian.
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map and let \(H\in M_{D}(\mathbb {C})\) be Hermitian, with Jordan decompositions \(H=P_+-P_-\) and \(T(H)=Q_+-Q_-\). Then \(\operatorname{tr}[Q_+]\leq \operatorname{tr}[P_+]\).
The matrix \(T(H)\) is Hermitian by Lemma 21.1.15. Let \(\Pi _+\) be the support projection of \(Q_+\) from Lemma 21.1.13, so that \(\Pi _+T(H)=Q_+\) and \(0\leq \Pi _+\leq \mathbb {1}\). Then
where the first inequality uses \(\operatorname{tr}[\Pi _+T(P_-)]\geq 0\) and the second uses \(\Pi _+\leq \mathbb {1}\) together with \(T(P_+)\geq 0\) (Lemma 21.1.14); the final step is trace preservation.
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map. Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\lVert T(H)\rVert _{\operatorname{tr}}\leq \lVert H\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Theorem 8.16 ] .
Write \(H=P_+-P_-\) and \(T(H)=Q_+-Q_-\) for the Jordan decompositions. Applying Lemma 21.1.16 to \(H\) and to \(-H\) gives \(\operatorname{tr}[Q_+]\leq \operatorname{tr}[P_+]\) and \(\operatorname{tr}[Q_-]\leq \operatorname{tr}[P_-]\), so by the Jordan trace-norm formula (Lemma 21.1.12),
Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map. Then for all density matrices \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\), \(\lVert T(\rho _1)-T(\rho _2)\rVert _{\operatorname{tr}} \leq \lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Eq. (8.80) ] .
By linearity \(T(\rho _1)-T(\rho _2)=T(\rho _1-\rho _2)\), and \(\rho _1-\rho _2\) is Hermitian as a difference of positive semidefinite matrices, so Theorem 21.1.17 applies.
21.2 Von Neumann entropy
For a Hermitian matrix \(\rho \in M_{D}(\mathbb {C})\) with eigenvalues \(\lambda _0,\ldots ,\lambda _{D-1}\), the von Neumann entropy is
where \(0\log 0:=0\).
If two Hermitian matrices are equal, then their von Neumann entropies are equal.
Substitute the equality of the matrices in the defining eigenvalue sum.
The zero matrix has zero von Neumann entropy: \(S(0)=0\).
All eigenvalues of the zero matrix are zero, and \(0\log 0=0\).
For any density matrix \(\rho \), \(S(\rho )\ge 0\).
Each eigenvalue \(\lambda _i\) of a density matrix satisfies \(0\le \lambda _i\le 1\), and \(-x\log x\ge 0\) on \([0,1]\).
The eigenvalues of a density matrix sum to \(1\).
Follows from \(\operatorname{tr}(\rho )=\sum _i\lambda _i=1\).
Each eigenvalue of a density matrix lies in \([0,1]\).
Non-negativity comes from positive semidefiniteness. The upper bound follows because the eigenvalues are non-negative and sum to \(1\).
For a density matrix \(\rho \in M_{D}(\mathbb {C})\) with \(D\ge 1\), one has \(S(\rho )\le \log D\).
By Jensen’s inequality applied to the concave function \(-x\log x\), the entropy is maximized when all eigenvalues equal \(1/D\).
For a density matrix \(\rho \) of rank \(r\), one has \(S(\rho )\le \log r\). This refines the dimension bound: only the nonzero eigenvalues contribute to the entropy, and there are exactly \(r\) of them.
The entropy is the \(-x\log x\) sum over all eigenvalues; the \(D-r\) zero eigenvalues contribute nothing. Jensen’s inequality applied to \(-x\log x\) over the \(r\) nonzero eigenvalues, with uniform weights \(1/r\), gives \(S(\rho )\le \log r\), the maximum attained when each nonzero eigenvalue equals \(1/r\). The rank equals the number of nonzero eigenvalues of the Hermitian matrix \(\rho \).
The von Neumann entropy of a Hermitian matrix is the \(-x\log x\) sum over the real parts of the roots of its characteristic polynomial \(\chi _\rho \):
The eigenvalues of a Hermitian matrix are exactly the roots of its characteristic polynomial, counted with multiplicity, so the eigenvalue sum defining \(S(\rho )\) equals the displayed sum over roots.
Let \(\rho \) be a Hermitian matrix and let \(\log \rho \) be its logarithm defined through the functional calculus. Then
A Hermitian matrix and its logarithm are simultaneously diagonalized by a unitary \(U\) with \(\rho =U\operatorname{diag}(\lambda _i)U^*\), so
Its negative is \(\sum _i({-}\lambda _i\log \lambda _i)=S(\rho )\). Zero eigenvalues contribute nothing under the convention \(0\log 0=0\), so no full-support assumption is required.
The logarithm is the totalized real logarithm, with \(\log x=\log \lvert x\rvert \) and \(\log 0=0\); both sides of the identity use it on every eigenvalue, so the equality holds for an arbitrary Hermitian matrix. It coincides with the physical entropy \(-\operatorname{tr}(\rho \log \rho )\) precisely when \(\rho \) is positive semidefinite. In that case (a density matrix \(\rho \), positive semidefinite with unit trace) the eigenvalues \(\lambda _i\) are non-negative, \(\rho \log \rho \) is Hermitian with real trace, and the real-part extraction is superfluous: the identity reduces to the standard expression \(S(\rho )=-\operatorname{tr}(\rho \log \rho )\).
For matrices \(\rho ,\sigma \in M_{D}(\mathbb {C})\), define the trace-log expression
On the physical domain where \(\rho \) is a density matrix and \(\sigma \) is positive definite, this is the Umegaki relative entropy.
For matrices \(\rho ,\sigma \in M_{D}(\mathbb {C})\),
Expand the matrix product over the difference \(\log \rho -\log \sigma \) and use linearity of the trace and of the real part.
For every matrix \(\rho \in M_{D}(\mathbb {C})\), one has \(D(\rho \Vert \rho )=0\).
The logarithmic difference \(\log \rho -\log \rho \) vanishes.
For every matrix \(\sigma \in M_{D}(\mathbb {C})\), one has \(D(0\Vert \sigma )=0\).
The trace-log expression is multiplied on the left by the zero matrix.
If \(\rho \) is Hermitian, then
Apply Lemma 21.2.13 to write
The trace-logarithm identity \(\operatorname{Re}\operatorname{tr}(\rho \log \rho )=-S(\rho )\) then gives the result.
Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices with \(\sigma \) of full rank. Then the relative entropy is non-negative, \(D(\rho \Vert \sigma )\ge 0\).
Diagonalize \(\rho =\sum _i p_i|e_i\rangle \! \langle e_i|\) and \(\sigma =\sum _j q_j|f_j\rangle \! \langle f_j|\) in their eigenbases, with all \(q_j{\gt}0\) since \(\sigma \) has full rank. The overlap numbers \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\) are non-negative with row sums and column sums equal to \(1\), because the two eigenbases are orthonormal. A trace computation gives
The row and column sums, together with the trace-one normalizations, give \(\sum _{i,j}P_{ij}(p_i-q_j)=\sum _i p_i-\sum _jq_j=0\). Hence
Each bracket is non-negative. If \(p_i{\gt}0\), put \(x=q_j/p_i{\gt}0\); the bracket is \(p_i(x-1-\log x)\geq 0\) by \(\log x\leq x-1\). If \(p_i=0\), the bracket equals \(q_j{\gt}0\). Therefore \(D(\rho \Vert \sigma )\geq 0\).
Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices satisfying the support condition \(\ker \sigma \subseteq \ker \rho \), that is, every vector annihilated by \(\sigma \) is annihilated by \(\rho \). Then the relative entropy is non-negative, \(D(\rho \Vert \sigma )\ge 0\).
Regularize \(\sigma \) by the trace-one perturbation \(\sigma _\varepsilon '=(1+\varepsilon D)^{-1}(\sigma +\varepsilon \mathbb {1})\), which is positive definite for every \(\varepsilon {\gt}0\), hence of full rank. The full-rank Klein inequality (Theorem 21.2.17) gives \(D(\rho \Vert \sigma _\varepsilon ')\ge 0\). The perturbation shares the eigenbasis \(\{ |f_j\rangle \} \) of \(\sigma \), so the cross term reduces to a scalar sum over the eigenvalues \(q_j\) of \(\sigma \),
For \(q_j{\gt}0\) the scalar logarithm converges to \(\log q_j\); for \(q_j=0\) the eigenvector \(|f_j\rangle \) lies in \(\ker \sigma \subseteq \ker \rho \), so its diagonal weight \(\langle f_j|\rho |f_j\rangle \) vanishes and the summand is identically zero. No eigenvalue-continuity input is needed, so \(D(\rho \Vert \sigma _\varepsilon ')\to D(\rho \Vert \sigma )\) as \(\varepsilon \to 0^+\) and the inequality passes to the limit.
Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices with \(\sigma \) of full rank. Then the relative entropy vanishes exactly when the states coincide, \(D(\rho \Vert \sigma )=0\iff \rho =\sigma \). Together with nonnegativity, this is the order property that makes \(D\) a divergence.
If \(\rho =\sigma \) then \(\log \rho -\log \sigma =0\) and \(D(\rho \Vert \sigma )=0\). For the converse, diagonalize \(\rho =\sum _i p_i|e_i\rangle \! \langle e_i|\) and \(\sigma =\sum _j q_j|f_j\rangle \! \langle f_j|\), with all \(q_j{\gt}0\) by full rank, and set \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\). As in Theorem 21.2.17,
a sum of non-negative terms: each is bounded below by the tangent inequality \(\log x\le x-1\) at \(x=q_j/p_i\), while the linear remainder telescopes through the doubly stochastic row and column sums, \(\sum _{i,j}P_{ij}(p_i-q_j)=\sum _ip_i-\sum _jq_j=0\). If the total vanishes then every term vanishes. On a row with \(p_i=0\) the term reads \(P_{ij}q_j\), so \(P_{ij}=0\); on a row with \(p_i{\gt}0\) the term forces the tangent inequality to be tight, \(\log (q_j/p_i)=q_j/p_i-1\), and strict concavity of the logarithm (\(\log x{\lt}x-1\) for \(x\neq 1\)) gives \(q_j=p_i\). Hence \(q_j=p_i\) whenever \(\langle e_i | f_j \rangle \neq 0\). Writing \(W=U_\rho ^\dagger U_\sigma \) for the overlap of the eigenvector unitaries, this matching says \(W\operatorname{diag}(q)=\operatorname{diag}(p)W\). Thus, conjugating \(\sigma \) into the eigenbasis of \(\rho \) and using the unitarity \(WW^\dagger =1\),
the spectral diagonal of \(\rho \). Therefore \(\sigma =U_\rho \operatorname{diag}(p)U_\rho ^\dagger =\rho \).
On pairs of positive definite matrices in \(M_{D}(\mathbb {C})\), the map \((\rho ,\sigma )\mapsto D(\rho \Vert \sigma )\) is jointly convex.
For \(s\in [0,1)\) consider the approximant
The trace \(\operatorname{tr}\rho \) is real-affine in the pair, hence convex, while \((\rho ,\sigma )\mapsto \operatorname{Re}\operatorname{tr}(\rho ^s\sigma ^{1-s})\) is jointly concave by the \(K=\mathbb {1}\) case of the Lieb concavity theorem (Corollary 20.6.14). Their difference, scaled by the non-negative factor \((1-s)^{-1}\), is therefore jointly convex, so each \(g_s\) is convex.
Writing \(\rho \) and \(\sigma \) in their eigenbases with eigenvalues \(p_i,q_j{\gt}0\) and overlap weights \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\), the approximant becomes the double sum \(g_s(\rho ,\sigma )=\sum _{i,j}(1-s)^{-1}(p_i-p_i^sq_j^{1-s})P_{ij}\). As \(s\to 1^-\) each per-pair term converges, via the scalar limit \((c^u-1)/u\to \log c\), to \(p_i(\log p_i-\log q_j)P_{ij}\), whose sum equals \(D(\rho \Vert \sigma )\). Thus \(g_s\to D\) pointwise on positive definite pairs, and since the pointwise limit of convex functions is convex, \(D\) is jointly convex.
On pairs of density matrices \((\rho ,\sigma )\) in \(M_{D}(\mathbb {C})\) satisfying the support condition \(\ker \sigma \subseteq \ker \rho \), the map \((\rho ,\sigma )\mapsto D(\rho \Vert \sigma )\) is jointly convex.
The domain is convex: for a strict convex combination of two such pairs the kernel of \(a\sigma _1+b\sigma _2\) is \(\ker \sigma _1\cap \ker \sigma _2\), since the quadratic forms of the positive semidefinite summands are non-negative and their positively weighted sum vanishes only when each does, and a positive semidefinite matrix annihilates exactly the vectors of zero quadratic form; the two pointwise support inclusions then give \(\ker \sigma _1\cap \ker \sigma _2\subseteq \ker \rho _1\cap \ker \rho _2\).
Regularize both arguments through the affine trace-one perturbation \(M_\varepsilon =(1+\varepsilon N)^{-1}(M+\varepsilon \mathbb {1})\), where \(N\) is the matrix dimension, which is positive definite for every \(\varepsilon {\gt}0\). Because the perturbation is affine, it commutes with the convex combination, so the four regularized endpoints and the regularized mixture all lie in the positive definite domain and the positive definite joint convexity (Theorem 21.2.20) gives the two-point inequality for the regularized pairs. As \(\varepsilon \to 0^+\), the regularized relative entropy converges on each pair of the support domain: \(D(\rho _\varepsilon \Vert \sigma _\varepsilon ) \to D(\rho \Vert \sigma )\). Indeed the perturbation shares the eigenbasis of its argument, so each trace-logarithm term is the diagonal sum \(\sum _jw_j(\varepsilon ) \log ((1+\varepsilon N)^{-1}(q_j+\varepsilon ))\), where \(q_j\) runs over the eigenvalues of \(\sigma \) and \(w_j(\varepsilon )\) is the corresponding diagonal weight. For \(q_j{\gt}0\), the scalar factor converges to \(\log q_j\), while at a zero eigenvalue the support condition makes the weight vanish, with \(w_j(\varepsilon )=(1+\varepsilon N)^{-1}\varepsilon \), so the summand is
by \(x\log x\to 0\) as \(x\to 0^+\). The two-point inequality therefore passes to the limit, giving joint convexity on the support domain.
For a Hermitian matrix \(A\), a real function \(f\), and a unitary \(U\), the continuous functional calculus satisfies
The conjugation \(\varphi :x\mapsto UxU^\dagger \) is a star-algebra automorphism of the matrix algebra, and the continuous functional calculus commutes with such automorphisms, \(\varphi (f(A))=f(\varphi (A))\); since \(\varphi (A)=UAU^\dagger \), this is the claim.
For a Hermitian matrix \(A\) and a unitary \(U\),
This is the \(f=\log \) case of Lemma 21.2.22: \(f(UAU^\dagger )=Uf(A)U^\dagger \) specialized to the real logarithm gives \(\log (UAU^\dagger )=U(\log A)U^\dagger \).
For Hermitian matrices \(\rho ,\sigma \) and a unitary \(U\),
Carry the two logarithms through the conjugation by Lemma 21.2.23, so that
where the middle equality is trace cyclicity together with \(U^\dagger U=1\).
For positive definite matrices \(\rho \) and \(\tau \),
The tensor product factors as \(\rho \otimes \tau =(\rho \otimes \mathbb {1})(\mathbb {1}\otimes \tau )\) into a pair of commuting positive definite matrices, so the logarithm of the product is the sum of the logarithms of the factors. Each unital embedding \(A\mapsto A\otimes \mathbb {1}\) and \(B\mapsto \mathbb {1}\otimes B\) is a continuous star-algebra homomorphism, hence commutes with the functional calculus, which carries each logarithm onto its factor: \(\log (\rho \otimes \mathbb {1})=\log \rho \otimes \mathbb {1}\) and \(\log (\mathbb {1}\otimes \tau )=\mathbb {1}\otimes \log \tau \).
For positive definite matrices \(\rho ,\sigma \) and a positive definite matrix \(\tau \) of unit trace,
Splitting each tensor logarithm by Lemma 21.2.25, the common \(\mathbb {1}\otimes \log \tau \) terms cancel in the difference, leaving \(\log (\rho \otimes \tau )-\log (\sigma \otimes \tau ) =(\log \rho -\log \sigma )\otimes \mathbb {1}\). Hence
where the trace of a tensor product factors as a product of traces. As \(\operatorname{tr}\tau =1\), the right-hand side is \(D(\rho \Vert \sigma )\).
Let \(\zeta \) be a primitive \(d\)-th root of unity and let \(i,j\) range over the residues modulo \(d\). Then
Each summand equals \(\xi ^b\) with \(\xi =\zeta ^i\overline{\zeta }^{\, j}\), and \(\xi ^d=1\). When \(i=j\) the base \(\xi \) equals \(1\) and the sum is \(d\). When \(i\ne j\) the base is a root of unity different from \(1\), so the geometric sum \((\xi -1)\sum _b\xi ^b=\xi ^d-1=0\) forces the sum to vanish; the equivalence \(\xi =1\iff i=j\) uses that \(\overline{\zeta }=\zeta ^{-1}\) and the injectivity of \(b\mapsto \zeta ^b\) on residues.
Fix a dimension \(d\ge 1\) and a primitive \(d\)-th root of unity \(\zeta \). The cyclic shift \(X\), the clock operator \(Z\), and every Weyl operator \(W(a,b)=X^aZ^b\) are unitary.
The cyclic shift is the permutation matrix of a cyclic permutation. The clock operator is diagonal, and each diagonal entry is a power of \(\zeta \), hence has modulus one. Thus \(X\) and \(Z\) are unitary, and so is every product \(X^aZ^b\).
Fix a dimension \(d\ge 1\) and a primitive \(d\)-th root of unity \(\zeta \), and let \(X\) be the cyclic shift \(|i\rangle \mapsto |i+1\rangle \) and \(Z=\operatorname{diag}(\zeta ^0,\ldots ,\zeta ^{d-1})\) the clock operator. For every matrix \(M\) on \(\mathbb {C}^d\), the uniform average of the conjugations by the \(d^2\) Weyl operators \(W(a,b)=X^aZ^b\) is the completely depolarizing channel:
The double average factors into a clock average followed by a shift average. The clock average \(\sum _bZ^bM(Z^b)^\dagger \) multiplies the entry \(M_{ij}\) by \(\sum _b\zeta ^{bi}\overline{\zeta ^{bj}}\), which by Lemma 21.2.27 is \(d\) when \(i=j\) and \(0\) otherwise; the result is \(d\) times the diagonal part of \(M\). The shift average \(\sum _aX^a(\operatorname{diag}v)(X^a)^\dagger \) cyclically permutes the diagonal entries, so each diagonal position receives the full sum \(\sum _kv_k=\operatorname{tr}M\), giving \((\operatorname{tr}M)\mathbb {1}\). Combining the two factors of \(d\) with the prefactor \(d^{-2}\) leaves \((\operatorname{tr}M/d)\mathbb {1}\).
Fix a dimension \(d_C\ge 1\) and a primitive \(d_C\)-th root of unity \(\zeta \). For every matrix \(M\) on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\), the uniform average of the conjugations by the \(d_C^2\) unitaries \(\mathbb {1}_S\otimes W(a,b)\) on the second factor is the partial trace over that factor tensored with the maximally mixed state \(\mathbb {1}_C/d_C\):
On each pair of blocks indexed by the first factor, the conjugation by \(\mathbb {1}_S\otimes W(a,b)\) acts as the Weyl conjugation \(W(a,b)(\cdot )W(a,b)^\dagger \) of the corresponding block of \(M\). Averaging over the \(d_C^2\) Weyl operators sends each block to the depolarizing channel by Theorem 21.2.29, replacing it by its trace times \(\mathbb {1}_C/d_C\). The block trace is exactly the corresponding entry of the partial trace \(\operatorname{tr}_C M\), so the average is \((\operatorname{tr}_C M)\otimes (\mathbb {1}_C/d_C)\).
For \(d_C\ge 1\), the matrix \(\tau _C=d_C^{-1}\mathbb {1}_C\) is positive definite and has trace one.
The identity is positive definite and \(d_C^{-1}{\gt}0\), so \(\tau _C\) is positive definite. Moreover, \(\operatorname{tr}\tau _C=d_C^{-1}\operatorname{tr}\mathbb {1}_C=1\).
For Hermitian matrices \(\rho ,\sigma \) on a finite index set and any bijection \(e\) from that set onto another finite set,
Write \(D(\rho \Vert \sigma ) =\operatorname{Re}\operatorname{tr}(\rho (\log \rho -\log \sigma ))\). The matrix logarithm is covariant under reindexing, \(\log (\rho _{e^{-1},e^{-1}}) =(\log \rho )_{e^{-1},e^{-1}}\), because reindexing is a star-algebra isomorphism and so commutes with the continuous functional calculus. The same isomorphism preserves products and the trace,
Taking \(M=\log \rho -\log \sigma \) and applying these three identities gives \(D(\rho _{e^{-1},e^{-1}}\Vert \sigma _{e^{-1},e^{-1}}) =D(\rho \Vert \sigma )\).
For positive definite matrices \(\rho ,\sigma \) on a tensor product of a system factor and an ancilla factor,
where \(\operatorname{tr}_C\) is the partial trace over the ancilla factor. This is the positive-definite base case; the source inequality on the support domain \(\ker \sigma \subseteq \ker \rho \) is Theorem 21.2.34.
By ancilla additivity (Theorem 21.2.26) the reduced-state relative entropy equals \(D\bigl((\operatorname{tr}_C\rho )\otimes (\mathbb {1}_C/d_C)\Vert (\operatorname{tr}_C\sigma )\otimes (\mathbb {1}_C/d_C)\bigr)\). By Lemma 21.2.30 each tensored reduced state is the uniform average of the conjugations by the \(d_C^2\) unitaries \(U_{ab}=\mathbb {1}_S\otimes W(a,b)\), so this is the relative entropy of a convex combination of the conjugated pairs \((U_{ab}\rho U_{ab}^\dagger ,U_{ab}\sigma U_{ab}^\dagger )\) with equal weights \(d_C^{-2}\). Joint convexity bounds it above by the same convex combination of the per-term relative entropies, each of which equals \(D(\rho \Vert \sigma )\) by unitary invariance (Theorem 21.2.24). As the weights sum to one, the bound is \(D(\rho \Vert \sigma )\).
For positive semidefinite \(\rho ,\sigma \) on a tensor product of a system factor and an ancilla factor of dimension \(d_C\), with the support condition \(\ker \sigma \subseteq \ker \rho \),
where \(\operatorname{tr}_C\) is the partial trace over the ancilla factor.
Let \(N\) be the dimension of the joint space. Regularize both arguments through the affine trace-shrinking perturbation \(M_\varepsilon =(1+N\varepsilon )^{-1}(M+\varepsilon \mathbb {1})\), which is positive definite for every \(\varepsilon {\gt}0\), so the positive-definite data-processing inequality (Theorem 21.2.33) gives
The right-hand side converges to \(D(\rho \Vert \sigma )\) as \(\varepsilon \to 0^+\), by the same shared-eigenbasis scalar-limit argument as the joint convexity on the support domain (Theorem 21.2.21).
For the left-hand side, the partial trace of the regularization is a differently scaled regularization of the marginal: because the partial trace is linear and \(\operatorname{tr}_C\mathbb {1}=d_C\mathbb {1}\),
This is the affine regularization of \(\operatorname{tr}_C M\) with the same scaling rate \(N\) but shift rate \(d_C\) and the smaller identity on the system factor. The support condition transfers to the marginals, \(\ker (\operatorname{tr}_C\sigma )\subseteq \ker (\operatorname{tr}_C\rho )\): a vector annihilated by \(\operatorname{tr}_C\sigma \) has vanishing marginal quadratic form, which splits into the non-negative joint quadratic forms of its single-ancilla lifts, so each lift lies in \(\ker \sigma \), hence in \(\ker \rho \), and reassembling the ancilla sum shows the vector lies in \(\ker (\operatorname{tr}_C\rho )\). With this support condition the arbitrary-rate affine regularization has the same both-arguments limit:
Passing the inequality through the two limits gives the support-domain bound.
Let \(\tau =\sum _i\lambda _i|i\rangle \! \langle i|\) be positive semidefinite. Its inverse square root on the support is
The inverse square root of a positive semidefinite matrix on its support is Hermitian.
The inverse square root of a positive semidefinite matrix on its support is positive semidefinite.
On the non-negative spectrum of \(\tau \), the defining function satisfies
The spectral functional calculus therefore gives \(f(\tau )\geq 0\).
If \(\tau \) is positive definite, then its inverse square root on the support is its ordinary inverse square root: \(\tau ^{-1/2}_{\mathrm{supp}}=(\sqrt\tau )^{-1}\).
Every eigenvalue of \(\tau \) is strictly positive. Hence the function defining the support inverse square root agrees on the spectrum with the reciprocal of the positive square-root function.
If \(P_\tau \) is the orthogonal projector onto the support of a positive semidefinite matrix \(\tau \), then
In an eigenbasis of \(\tau \), the left-hand side has eigenvalue zero when \(\lambda _i=0\) and eigenvalue \(\lambda _i^{-1/2}\lambda _i\lambda _i^{-1/2}=1\) otherwise.
If \(P_\tau \) is the orthogonal projector onto the support of a positive semidefinite matrix \(\tau \), then
The support inverse square root commutes with \(\tau \) by spectral functional calculus, so
where the last step is Lemma 21.2.39. Taking adjoints gives the identity with \(\tau \) on the left.
Let \(A\) and \(B\) be positive semidefinite, and let \(f\colon \mathbb {R}\to \mathbb {R}\) be multiplicative on the non-negative reals. Then
Diagonalize \(A\) and \(B\). In the resulting product eigenbasis, the eigenvalues of \(A\otimes B\) are \(a_i b_j\) with \(a_i,b_j\geq 0\), and the claim follows from \(f(a_i b_j)=f(a_i)f(b_j)\).
Let \(A\) and \(B\) be positive semidefinite. Their positive square roots, support inverse square roots, and support projections satisfy
Moreover,
If \(A\) is positive definite, then \(P_A=\mathbf1\).
Diagonalize both factors. The eigenvalues of \(A\otimes B\) are the products \(a_i b_j\). Both the square-root function and the function that equals \(x^{-1/2}\) for \(x{\gt}0\) and zero at \(x=0\) are multiplicative on the non-negative reals. For the support projectors, apply the sandwich identity to \(A\otimes B\), factor the support inverse square root and matrix products, and apply the sandwich identity to each factor. The cancellation identities follow entrywise. If \(A\) is positive definite, then \(A^{-1/2}_{\mathrm{supp}}=(\sqrt A)^{-1}\) and \(\sqrt A\) is invertible. Hence the cancellation identity gives \(P_A=\sqrt A(\sqrt A)^{-1}=\mathbf1\).
Let \(\tau \) be positive semidefinite and let \(c{\gt}0\). Then
For every finite-dimensional auxiliary space \(\mathcal H_R\),
The corresponding identity for an auxiliary left factor is
Consequently, for \(d_R{\gt}0\),
The scalar identity follows from the functional calculus and \(\sqrt{cx}=\sqrt c\sqrt x\) for \(c{\gt}0\). For the support-inverse function \(f(x)=x^{-1/2}\) when \(x{\gt}0\) and \(f(0)=0\), the functional calculus gives
This proves (??); the same argument with the identity as the left factor proves (??). Applying (??) with \(c=d_R^{-1}\) to \(\tau \otimes \mathbf1_R\) then gives (??).
Let \(\sigma \) be positive semidefinite on \(H_L\otimes H_R\), and set \(\tau =\operatorname{tr}_R\sigma \). The Petz transpose formula on the support of \(\tau \) is
This is the support formula of [ HJPW04 , Theorem 3, equation (8) ] . It is not asserted to be trace preserving on operators outside the support of \(\tau \).
For every matrix \(X\), the support Petz map is given by (??).
Let \(\rho _A\) and \(\rho _{BC}\) be positive semidefinite, and define
We use the canonical reassociation from \(A\times (B\times C)\) to \((A\times B)\times C\), so that the right partial trace removes \(C\).
Set \(\rho _B=\operatorname{tr}_C\rho _{BC}\), and let \(P_A\) and \(P_B\) be the support projections of \(\rho _A\) and \(\rho _B\). Then
Each identity is understood after the same canonical reassociation of the three tensor factors.
Expand the marginal in a product basis. The remaining identities follow from functional calculus for positive semidefinite tensor products. The support projection is obtained by multiplying the factorized support inverse square root on both sides of the marginal.
If \(P_A\) is the support projection of \(\rho _A\), define
For the reference in (??), the raw Petz map for \(\operatorname{tr}_C\) satisfies
Equivalently, for every product operator \(A_0\otimes X_B\),
The formulas use the canonical identification \((H_A\otimes H_B)\otimes H_C\cong H_A\otimes (H_B\otimes H_C)\).
This is the globally valid ambient-space form of [ HJPW04 , equation (10) ] . The literal identity-tensored formula in that equation is obtained on the support of \(\rho _A\), or after choosing an extension away from that support. For singular \(\rho _A\), the raw map on the full matrix algebra contains the compression \(X\mapsto P_AXP_A\).
Substitute the tensor factorizations of the square root and marginal support inverse into the Petz sandwich. The first-factor terms reduce to \(P_AA_0P_A\). This proves the formula for product operators. A finite product-operator decomposition proves the linear-map identity.
Let \(X\) be an operator on \(H_A\otimes H_B\). If
then
If \(\rho _A\) is positive definite, then \(P_A=\mathbf1_A\), and this identity holds for every \(X\).
When \(\rho _A\) is singular, no global identity-tensored formula is asserted for the raw map outside the displayed support. Nor is the generic trace-preserving completion in Definition 21.2.61 asserted to factor: its complementary projection is \(\mathbf1_{AB}-P_A\otimes P_B\), which need not be the identity on \(A\) tensored with a projection on \(B\).
The support condition makes \(\mathcal S_{P_A}\) act as the identity on \(X\). If \(\rho _A\) is positive definite, its support projection is the identity, so the condition holds on the full matrix algebra.
Let \(d_A{\gt}0\) and let \(\rho _{BC}\) be positive semidefinite. Define
Under the canonical reassociation from \(A\times (B\times C)\) to \((A\times B)\times C\), the right partial trace removes \(C\).
For the reference in (??),
Expand the two partial traces in a product basis. The sum over the \(d_A\) diagonal entries cancels the factor \(d_A^{-1}\). Trace invariance under a partial trace then gives the last identity.
If \(\rho _{BC}\) is positive semidefinite and \(\rho _B=\operatorname{tr}_C\rho _{BC}\), then
The marginal identity gives \(\sigma _{AB}=d_A^{-1}\mathbf1_A\otimes \rho _B\). Apply the support-inverse scaling identity and the tensor identity for an auxiliary left factor. Since \(d_A{\gt}0\), \((\sqrt{d_A^{-1}})^{-1}=\sqrt{d_A}\).
For the reference in (??), the support Petz map for \(\operatorname{tr}_C\) factors as
after the canonical identification \((H_A\otimes H_B)\otimes H_C\cong H_A\otimes (H_B\otimes H_C)\). This is the maximally mixed specialization of [ HJPW04 , equation (10) ] . It is the tensor-product identity for the raw Petz support formula, not the Hayashi–Koashi–Imoto block decomposition.
Functional calculus gives \(\sqrt{\sigma _{ABC}} =d_A^{-1/2}\mathbf1_A\otimes \sqrt{\rho _{BC}}\). Theorem 21.2.53 gives the corresponding factorization of the marginal support inverse. The scalar factors cancel in the Petz sandwich. The identity follows first for \(A_0\otimes X_B\) and then for every operator by a finite sum of product operators.
If \(P_B\) is the support projector of \(\rho _B\), then the support projector of \(d_A^{-1}\mathbf1_A\otimes \rho _B\) is
Insert the factorized support inverse into \(P_{AB}=\sigma _{AB,\mathrm{supp}}^{-1/2} \sigma _{AB}\sigma _{AB,\mathrm{supp}}^{-1/2}\). The scalar factors cancel, and the remaining sandwich is \(\mathbf1_A\otimes (\rho _{B,\mathrm{supp}}^{-1/2}\rho _B \rho _{B,\mathrm{supp}}^{-1/2})=\mathbf1_A\otimes P_B\).
The support Petz transpose map \(\mathcal R_\sigma \) is completely positive.
For an orthonormal basis \((e_r)_r\) of \(H_R\), let \(J_r:H_L\to H_L\otimes H_R\) be given by \(J_r(v)=v\otimes e_r\). Then
Thus the support Petz map has a rectangular Kraus representation.
Let \(P_\tau \) be the orthogonal projector onto the support of \(\tau =\operatorname{tr}_R\sigma \). Then, for every matrix \(X\),
Cyclicity of the trace and the defining property of the partial trace give
By (??), this equals \(\operatorname{tr}(P_\tau X)\).
Put \(Q_\tau =\mathbf1_L-P_\tau \) and \(\omega _R=\operatorname{tr}_L\sigma \). The complementary term is
The complementary preparation term \(\mathcal C_\sigma \) is completely positive.
The map \(X\mapsto Q_\tau XQ_\tau \) has the single Kraus operator \(Q_\tau \). Adjoining the positive semidefinite matrix \(\omega _R\) is completely positive, and the composition of these two maps is completely positive.
If \(\sigma \) has trace one, then
Since \(\operatorname{tr}(\omega _R)=\operatorname{tr}(\sigma )=1\) and \(Q_\tau ^2=Q_\tau \), factorization and cyclicity of the trace give \(\operatorname{tr}(\mathcal C_\sigma (X)) =\operatorname{tr}(Q_\tau XQ_\tau ) =\operatorname{tr}(Q_\tau ^2X) =\operatorname{tr}(Q_\tau X)\).
The completed Petz map is
If \(\sigma \) is positive semidefinite with trace one, then \(\widehat{\mathcal R}_\sigma \) is completely positive and trace preserving.
The two summands are completely positive. Since \(\operatorname{tr}(\omega _R)=1\), the complementary term satisfies (??). Adding (??) to (??) gives trace preservation.
For the reference in (??), the chosen complementary preparation term factors as
after the canonical reassociation of the three tensor factors. This follows from the chosen support completion in (??), not from [ HJPW04 , equation (10) ] .
From (??), \(Q_{AB}=\mathbf1_A\otimes Q_B\), while \(\operatorname{tr}_{AB}\sigma _{ABC}=\rho _C\). Hence
A finite product-operator decomposition gives the result for every input.
For the reference in (??),
after canonical reassociation of the three tensor factors. The raw support-map summand is the maximally mixed specialization of [ HJPW04 , equation (10) ] . The complementary summand comes from the chosen support completion in (??). This theorem does not assert a Hayashi–Koashi–Imoto decomposition.
Add (??) and (??), and use additivity of the tensor product of linear maps.
If \(\rho _{BC}\) is positive semidefinite with trace one, then \(\widehat{\mathcal R}_{\sigma _{ABC}}\) is completely positive and trace preserving.
By (??), \(\operatorname{tr}(\sigma _{ABC})=\operatorname{tr}(\rho _{BC})=1\). Apply the channel property of the completed Petz map.
If \(P_\tau XP_\tau =X\), then \(\widehat{\mathcal R}_\sigma (X)=\mathcal R_\sigma (X)\).
The assumption implies \(Q_\tau XQ_\tau =0\). Hence \(\mathcal C_\sigma (X)=0\), and adding the complementary term to \(\mathcal R_\sigma (X)\) does not change its value.
Let \(A\) be Hermitian, with spectral decomposition \(A=U\operatorname{diag}(\lambda _i)U^\dagger \). Its support projection is
Let \(A\) and \(B\) be Hermitian matrices such that \(\ker A\subseteq \ker B\). If \(P_A\) is the support projection of \(A\), then \(P_A B P_A=B\).
The complementary projection \(\mathbf1-P_A\) has range contained in \(\ker A\), and hence in \(\ker B\). Thus \(B(\mathbf1-P_A)=0\). Taking adjoints gives \((\mathbf1-P_A)B=0\), so \(B=P_A B\), and therefore \(P_A B P_A=P_A B=B\).
For \(w\in H_L\) and a distinguished basis vector \(e_r\in H_R\), define the lift \(J_r w=w\otimes e_r\in H_L\otimes H_R\).
Let \(X\) be a matrix on \(H_L\otimes H_R\). For basis indices \(i,s\) and \(r\),
Since \((J_r w)_{(j,c)}=w_j\mathbf1_{c=r}\), expansion of the matrix-vector product gives
Let \(X\) be a matrix on \(H_L\otimes H_R\), let \(w\in H_L\), and let \((e_r)_r\) be the distinguished orthonormal basis of \(H_R\). Then
Expanding the matrix products and the partial trace gives \(\sum _{i,j,r}\overline{w_i}X_{(i,r),(j,r)}w_j\) on both sides.
Let \(\sigma \) be positive semidefinite on \(H_L\otimes H_R\), and let \(\rho \) be any matrix on the same space such that \(\ker \sigma \subseteq \ker \rho \). Then
If \(w\in \ker (\operatorname{tr}_R\sigma )\), then
Each summand is non-negative and therefore vanishes. Positive semidefiniteness shows that \(\sigma (w\otimes e_r)=0\) for every \(r\), so the joint kernel inclusion gives \(\rho (w\otimes e_r)=0\). Summing the diagonal components over \(r\) yields \((\operatorname{tr}_R\rho )w=0\).
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(H_L\otimes H_R\) such that \(\ker \sigma \subseteq \ker \rho \). Then
Kernel inclusion descends through the partial trace, giving \(\ker (\operatorname{tr}_R\sigma )\subseteq \ker (\operatorname{tr}_R\rho )\). Theorem 21.2.68 therefore yields \(P_\tau (\operatorname{tr}_R\rho )P_\tau =\operatorname{tr}_R\rho \), where \(\tau =\operatorname{tr}_R\sigma \). By Theorem 21.2.66,
For every positive semidefinite \(\sigma \),
By (??), the middle factor reduces to \(P_\tau \otimes \mathbf1_R\). Marginal-support absorption gives \((\mathbf1_{LR}-P_\tau \otimes \mathbf1_R)\sigma =0\), so
because that matrix is positive semidefinite of trace zero. Therefore \(\sqrt\sigma (P_\tau \otimes \mathbf1_R)\sqrt\sigma =\sigma \).
For every positive semidefinite \(\sigma \),
The marginal satisfies \(P_\tau \tau P_\tau =\tau \). Therefore the completed channel agrees with the support Petz map at \(\tau \), and Theorem 21.2.74 gives \(\widehat{\mathcal R}_\sigma (\tau )=\sigma \).
Let \(\rho \) be a Hermitian matrix indexed by a finite set \(J\), and let \(e : I \to J\) be a bijection from a finite set \(I\). The reindexed matrix on \(I\) with entries \(\rho _{e(i)\, e(j)}\) has the same von Neumann entropy as \(\rho \).
By Lemma 21.2.9 the entropy depends only on the characteristic polynomial. Reindexing conjugates \(\rho \) by a permutation matrix, so \(\chi _{(\rho _{e(i)\, e(j)})}=\chi _\rho \), and the entropies coincide.
Let \(A \in M_{m \times n}(\mathbb {C})\) and \(B \in M_{n \times m}(\mathbb {C})\). Then the charpoly-root entropy sum is invariant under the cyclic swap \(AB \mapsto BA\):
with roots counted with algebraic multiplicity.
The rectangular characteristic-polynomial identity gives \(X^n\chi _{AB}=X^m\chi _{BA}\). Thus the two characteristic polynomials have the same nonzero roots, with multiplicity; the additional zero roots contribute nothing because \(0\log 0=0\).
For matrices \(A \in M_{m \times n}(\mathbb {C})\) and \(B \in M_{n \times m}(\mathbb {C})\) such that \(AB\) and \(BA\) are Hermitian, \(S(AB)=S(BA)\).
For density matrices \(\omega \) and \(\tau \), \(S(\omega \otimes \tau )=S(\omega )+S(\tau )\).
The tensor product is unitarily conjugate to the diagonal of eigenvalue products \(\lambda _i\mu _j\). Using
and the unit eigenvalue sums collapses the double sum to \(S(\omega )+S(\tau )\).
For a density matrix \(\omega \) and a scalar \(c\), \(S(c\, \omega )=c\, S(\omega )-c\log c\).
The eigenvalues of \(c\, \omega \) are \(c\lambda _i\), so the entropy is \(\sum _i-(c\lambda _i)\log (c\lambda _i)\). The splitting identity
together with the unit eigenvalue sum \(\sum _i\lambda _i=1\) gives
the stated form.
For a family of Hermitian matrices \(M_j\), the block-diagonal direct sum satisfies
Each block diagonalizes by a unitary congruence, and the block-diagonal assembly of the block unitaries diagonalizes the direct sum, whose eigenvalue multiset is the disjoint union of the block eigenvalue multisets. Hence
Let \(A=\sum _j M_j\) be a finite sum of Hermitian matrices. Suppose there are operators \(P_j\) such that
Then \(S(A)=\sum _j S(M_j)\). The operators \(P_j\) need only resolve the support of \(A\); their sum need not be the identity on the ambient space. This is the support form of the direct-sum entropy identity used in [ CPGSV16 , Appendix C.2, lines 1760–1770 ] .
Write \(A=XY\), where \(X\) maps the direct sum of the labelled ambient spaces to the original space by the matrices \(M_j\), and \(Y\) maps back by the operators \(P_j\). The support and annihilation identities give \(YX=\bigoplus _jM_j\). Reversing the two rectangular factors preserves the nonzero eigenvalues, while any additional eigenvalues are zero. Entropy is therefore unchanged, and additivity on the block diagonal gives the result.
Let \(\omega _j\) be density matrices and let \(p_j\geq 0\). Then
In particular, when the \(p_j\) form a probability distribution, this is
This is the entropy identity used in [ CPGSV16 , Appendix C.2, lines 1760–1770 ] .
Entropy is additive over the orthogonal blocks. Applying the scaled-state formula to each block gives \(S(p_j\omega _j)=-p_j\log p_j+p_jS(\omega _j)\), and summing over \(j\) gives the result.
Suppose \(p_j{\gt}0\) and \(L_j\leq R_j\) for every \(j\). If \(\sum _jp_jL_j=\sum _jp_jR_j\), then \(L_j=R_j\) for every \(j\). This is the positivity argument applied to strong subadditivity in [ CPGSV16 , Appendix C.2, lines 1770–1780 ] .
Each number \(p_j(R_j-L_j)\) is non-negative, and their sum vanishes. Therefore every one vanishes. Since \(p_j{\gt}0\), it follows that \(R_j-L_j=0\).
21.3 Tripartite partial traces
For a tripartite matrix \(\rho _{ABC}\) on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B} \otimes \mathbb {C}^{d_C}\), the partial trace over \(A\) is
The partial trace over \(C\) is
The partial trace over \(A\) and \(C\) is
If \(\rho _{ABC}\) is Hermitian, then \(\operatorname{tr}_A(\rho _{ABC})\), \(\operatorname{tr}_C(\rho _{ABC})\), and \(\operatorname{tr}_{AC}(\rho _{ABC})\) are all Hermitian. The same holds for bipartite partial traces \(\operatorname{tr}_A(\rho _{AB})\) and \(\operatorname{tr}_B(\rho _{AB})\).
Follows from \(\overline{\rho _{ji}} = \rho _{ij}\) applied entry-wise inside the summation defining each partial trace.
21.4 Strong subadditivity
Let \(\rho \) be a density matrix on \(A \otimes R\) with reduced state \(\rho _R = \operatorname{tr}_A \rho \), and suppose the support condition \(\ker ((\mathbb {1}_A / d_A) \otimes \rho _R) \subseteq \ker \rho \) holds. Then
When \(\rho _R\) is singular the tensor logarithm of the reference does not split. Regularize the reduced state by the affine perturbation \(\rho _{R,\varepsilon } = (1 + d_R\varepsilon )^{-1}(\rho _R + \varepsilon \mathbb {1})\), positive definite for \(\varepsilon {\gt} 0\). The reference \((\mathbb {1}_A / d_A) \otimes \rho _{R,\varepsilon }\) is then a positive definite tensor product, whose logarithm splits, so the cross trace term evaluates to \(-\log d_A + \operatorname{Re}\operatorname{tr}(\rho _R\log \rho _{R,\varepsilon })\). As \(\varepsilon \to 0^+\) the support condition makes the zero eigenvalues of \(\rho _R\) contribute nothing, so
Since \(D(\rho \| \sigma ) = -S(\rho ) - \operatorname{Re}\operatorname{tr}(\rho \log \sigma )\), the evaluation follows.
Let \(\rho \) be a positive semidefinite operator on \(A \otimes R\) with reduced state \(\rho _R = \operatorname{tr}_A \rho \). Then the singular reference \((\mathbb {1}_A / d_A) \otimes \rho _R\) satisfies the support condition \(\ker ((\mathbb {1}_A / d_A) \otimes \rho _R) \subseteq \ker \rho \).
A vector \(v\) annihilated by \((\mathbb {1}_A / d_A) \otimes \rho _R\) is annihilated by \(\mathbb {1}_A \otimes \rho _R\), since \(\mathbb {1}_A / d_A\) is invertible. Let \(P\) be the orthogonal projection onto the range of \(\rho _R\). The complementary lift \(\mathbb {1}_A \otimes (\mathbb {1}- P)\) then fixes \(v\), while it annihilates \(\rho \) on the left because the reduced state of \(\rho \) on \(R\) is supported on the range of \(P\). Hence \(\rho v = 0\).
For every tripartite matrix \(\rho _{ABC}\),
For indices \((a_1,b_1)\) and \((a_2,b_2)\), the corresponding matrix entry is
This is the corresponding entry of \((\mathbb {1}_A/d_A)\otimes \rho _B\).
For a tripartite density matrix \(\rho _{ABC}\),
Read the inequality as one instance of data processing under the partial trace over \(C\), with the singular reference state \(\sigma _{ABC} = (\mathbb {1}_A / d_A) \otimes \rho _{BC}\). Being a density operator is not enough to place the pair \((\rho _{ABC}, \sigma _{ABC})\) in the relative-entropy domain, which is the kernel inclusion \(\ker \sigma _{ABC} \subseteq \ker \rho _{ABC}\); the marginal support lemma supplies it. Against this reference the relative entropy of each pair evaluates to an entropy difference:
For a singular reference the tensor logarithm does not split, so each evaluation regularizes the reduced state through the affine perturbation \(\rho _{R,\varepsilon } = (1 + d_R\varepsilon )^{-1}(\rho _R + \varepsilon \mathbb {1})\), which is positive definite, and passes to the limit
where the zero eigenvalues of \(\rho _R\) contribute nothing because the kernel inclusion makes the corresponding diagonal weights vanish. Data processing on the singular support domain under the partial trace over \(C\), which sends \(\rho _{ABC} \mapsto \rho _{AB}\) and \((\mathbb {1}_A / d_A) \otimes \rho _{BC} \mapsto (\mathbb {1}_A / d_A) \otimes \rho _B\), reads
Substituting (??) and (??) into (??) cancels the common \(\log d_A\) and rearranges to the claimed inequality.
For a positive definite tripartite density matrix \(\rho _{ABC}\),
This is the positive definite case of Theorem 21.4.4.
Read the inequality as one instance of data processing under the partial trace over \(C\), with reference state \(\sigma _{ABC} = (\mathbb {1}_A / d_A) \otimes \rho _{BC}\), which is positive definite because \(\rho _{BC}\) is. Against this reference the relative entropy of each pair evaluates to an entropy difference:
Each equality follows from the tensor logarithm split and the adjoint of the partial trace. Data processing under the partial trace over \(C\), which sends \(\rho _{ABC} \mapsto \rho _{AB}\) and \((\mathbb {1}_A / d_A) \otimes \rho _{BC} \mapsto (\mathbb {1}_A / d_A) \otimes \rho _B\), reads
Substituting (??) and (??) into (??) cancels the common \(\log d_A\) and rearranges to the claimed inequality.
A tripartite density matrix \(\rho _{ABC}\) satisfies SSA equality if
Let \(A\) and \(B\) be positive semidefinite, with \(P_A\) and \(P_B\) the orthogonal projections onto their respective ranges. Then
Here the logarithm is extended by zero on the kernel. This identity is the analytic justification for the singular product reference used below; it is not stated verbatim in [ HJPW04 ] .
Diagonalize \(A\) and \(B\). On an eigenvector with eigenvalues \(a,b\geq 0\), (??) becomes
If \(a,b{\gt}0\), this is the ordinary product identity for the logarithm. With the kernel convention in the statement, both sides vanish if either eigenvalue is zero.
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices such that \(\ker \sigma \subseteq \ker \rho \), and let \(\tau \) be positive semidefinite with \(\operatorname{tr}\tau =1\). Then
Write \(P_\rho \), \(P_\sigma \), and \(P_\tau \) for the support projections. The kernel inclusion gives \(P_\sigma \rho P_\sigma =\rho \) and \(\rho P_\sigma =\rho \), while \(\rho P_\rho =\rho \) and \(\tau P_\tau =\tau \). Substituting the singular tensor-logarithm formulas into the relative entropy and using these four identities cancels the two contributions containing \(\log \tau \). Factoring the trace of the remaining tensor product gives
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\) such that \(\ker \sigma \subseteq \ker \rho \), and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For a primitive \(d_C\)-th root of unity, put \(U_{ab}=\mathbb {1}_S\otimes W(a,b)\) and
Then, for every \(a,b\),
This is a scalar equality-propagation prerequisite for [ HJPW04 , Theorem 3 and equation (8) ] ; it neither characterizes equality in joint convexity nor asserts recovery.
Let \(\tau _C=d_C^{-1}\mathbb {1}_C\). The twirl identity gives \(\overline X=(\operatorname{tr}_C X)\otimes \tau _C\). The support inclusion passes to the partial traces:
Hence support-domain ancilla additivity and the saturation hypothesis give
Every \(U_{ab}\) is unitary, and unitary invariance gives
Together, (??) and (??) prove (??).
Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\) such that \(\ker \sigma \subseteq \ker \rho \), and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For a primitive \(d_C\)-th root of unity, put \(U_{ce}=\mathbb {1}_S\otimes W(c,e)\) and
Then
This scalar identity is associated with the finite Jensen step in the Weyl proof of data processing. It is a prerequisite for [ HJPW04 , Theorem 3 and equation (8) ] ; it neither characterizes equality in joint convexity nor asserts recovery.
Unitary invariance gives, for every \(c,e\),
Theorem 21.4.9 gives \(D(\overline\rho \Vert \overline\sigma )=D(\rho \Vert \sigma )\). Substituting into the right-hand side gives
If \(A\) and \(B\) are positive definite and \(c{\gt}0\), then
The functional-calculus identity \(\log (cA)=(\log c)\mathbb {1}+\log A\), and its analog for \(B\), show that the scalar logarithmic terms cancel in \(\log (cA)-\log (cB)\). Linearity of the trace then gives the result.
Let \(\rho \) and \(\sigma \) be positive definite, and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For the uniformly weighted Weyl conjugates \(A_g=d_C^{-2}U_g\rho U_g^\dagger \) and \(B_g=d_C^{-2}U_g\sigma U_g^\dagger \), put \(A=\sum _gA_g\) and \(B=\sum _gB_g\). Then
Let \(a,b{\gt}0\). Then the function
is integrable on \((0,\infty )\), and
This is the scalar normalization \((\mathrm{intspec})\) in Jenčová–Ruskai, arXiv:0903.2895v4, §2.1, lines 406–413.
The integrand is the derivative of \(a(\log (1+t)-\log (a+tb))\). Its value at \(t=0\) is \(-a\log a\), whereas its limit as \(t\to \infty \) is \(-a\log b\). The derivative has constant sign, according as \(a-b\) is positive or negative, and is therefore integrable. The fundamental theorem of calculus gives the stated value.
Let \(A\) and \(B\) be Hermitian matrices of the same size, with spectral resolutions
Write \(w_{ij}=\langle u_i,v_j\rangle \), and use the total real logarithm: \(\log 0=0\), while \(\log x=\log |x|\) for \(x{\lt}0\). Then
This is an algebraic totalized extension to arbitrary Hermitian matrices of the homogeneous trace-log identity \((\mathrm{J1})\), which Jenčová–Ruskai state for strictly positive matrices in arXiv:0903.2895v4, lines 277–287.
Expand both trace terms in eigenbases. Unitarity of the overlap matrix gives \(\sum _j|w_{ij}|^2=1\), so
Subtraction gives (??).
Let \(A\) and \(B\) be positive semidefinite matrices of the same size and suppose that \(\ker B\subseteq \ker A\). With the spectral notation of Theorem 21.4.14, define, for \(t{\gt}0\),
Thus no ordinary quotient with \(\alpha _i=\beta _j=0\) is used. Define also
The function \(I_{A,B}\) is integrable on \((0,\infty )\), and
The trace-log identity \((\mathrm{J1})\), the scalar normalization \((\mathrm{intspec})\), and its matrix form \((\mathrm{intAB})\) occur in Jenčová–Ruskai, arXiv:0903.2895v4, at lines 277–287, 406–413, and 423–427, respectively. The support-domain extension is given at lines 717–720.
This theorem concerns the spectral expression \(I_{A,B}\). The next theorem identifies it with the coordinate-free left-right quadratic form underlying the finite Weyl formula.
If \(\beta _j=0\), the kernel inclusion gives \(\alpha _i|w_{ij}|^2=0\), and both \(r_{ij}\) and \(e_{ij}\) are defined to be zero. If \(\beta _j{\gt}0\), the scalar integral applies when \(\alpha _i{\gt}0\), while the term is identically zero when \(\alpha _i=0\). Since the double sum is finite, integration term by term gives
The kernel inclusion also shows that replacing each \(\beta _j=0\) summand in Theorem 21.4.14 by zero does not change its value. Hence that theorem identifies the sum in (??) with \(D(A\Vert B)\).
Let \(A\) and \(B\) be positive semidefinite matrices of the same size, let \(P_B\) be the orthogonal projection onto the support of \(B\), and, for \(t{\gt}0\), set
Write \(S_t^+\) for the generalized inverse that vanishes on \(\ker S_t\). Define
With the spectral notation of Theorem 21.4.15,
Consequently,
The function in (??) is continuous on \((0,\infty )\). If \(\ker B\subseteq \ker A\), then \(A P_B=A\), so \(Q_{A P_B}(t)\) equals the quadratic form with source \(\operatorname{vec}(A^{\top })\).
This is the support-projected form of \((\mathrm{intAB})\) in Jenčová–Ruskai, arXiv:0903.2895v4, §2.1, lines 423–427, with the support convention at lines 717–720.
Diagonalize \(A\) and \(B\) and write \(W=U_A^\ast U_B\). In these coordinates, the two equations
have entries
For every positive semidefinite \(S\), the identities \(S^+S=SS^+=P_S\) imply that \(Sx=b\) gives \(\langle b,S^+b\rangle =\langle b,x\rangle \). Applying this identity to (??) and (??), with (??) and (??), gives (??) and (??). Their coefficientwise combination is the scalar function appearing in Lemma 21.4.13. Continuity follows term by term from the finite sums.
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size. For \(t{\gt}0\), write
Then
No kernel inclusion between \(A_i\) and \(B_i\) is required. This lemma is a positive-semidefinite support-domain extension of the positive-definite calculation in equations \((\mathrm{Mj})\), \((\mathrm{eq:Schz1})\), and \((\mathrm{eq:Schwzt})\) at lines 1299–1328 of Jenčová–Ruskai, arXiv:0903.2895v4; their generalized-inverse notation is given at lines 254–262. The paper does not state this extension. Its later singular equality theorem at lines 761–785 assumes \(\ker B_i\subseteq \ker A_i\) and is not asserted here. The subsequent singular entropy-equality passage is recorded in docs/paper-gaps/cpsv16_ssa_equality_hayashi_markov.tex.
Put
The source equation for the support relative-modular operator shows that \(b_i\) lies in the support of \(S_i\). The same argument applied to \(\sum _i A_i\) and \(\sum _i B_i\) shows that \(\sum _i b_i\) lies in the support of \(\sum _i S_i\). The support-resolvent residual identity writes the displayed defect as a sum of nonnegative quadratic residuals.
Let \(A\) and \(B\) be positive definite, with spectral resolutions \(A=\sum _i\alpha _i|u_i\rangle \! \langle u_i|\) and \(B=\sum _j\beta _j|v_j\rangle \! \langle v_j|\). Put \(w_{ij}=\langle u_i,v_j\rangle \). Let \(L_A\) and \(R_B\) denote left and right multiplication, \(L_A(X)=AX\) and \(R_B(X)=XB\). Then, for \(t{\gt}0\),
Both quadratic forms are continuous on \((0,\infty )\). Moreover,
and the integrand in (??) is continuous and integrable on the positive half-line. This is the positive-definite spectral route of Jenčová–Ruskai, arXiv:0903.2895v4, §4.
Vectorization sends \(L_A+tR_B\) to \(A\otimes \mathbb {1}+t\mathbb {1}\otimes B^{\mathsf T}\). The vectors \(u_i\otimes \overline{v_j}\) diagonalize this matrix with eigenvalues \(\alpha _i+t\beta _j\), which proves (??) and (??). Insert these identities into (??), use \(\sum _i|w_{ij}|^2=1\), and apply Lemma 21.4.13 term by term. The same finite spectral sum proves continuity and integrability.
Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(S_i\) be a positive-semidefinite matrix, let \(P_i\) be its support projection, and let \(b_i\) lie in its support. Write
Assume also that \(b\) lies in the support of \(S\), and put \(x=Gb\). Then
Jenčová and Ruskai give the positive-definite residual expansion in equations \((\mathrm{Mj})\) and \((\mathrm{eq:Schz1})\) of the Appendix to arXiv:0903.2895v4. The support assumptions make the same expansion valid for the generalized inverses of the \(S_i\) and of \(S\).
Since \(G_iS_i=S_iG_i=P_i\) and \(P_i b_i=b_i\), expansion of the \(i\)th summand gives
Sum over \(i\). The support assumption on \(b\) gives \(Sx=SGb=b\), so the last three terms combine to \(-\langle b,Gb\rangle \).
Under the hypotheses and notation of Lemma 21.4.19, suppose that
Then, for every \(i\in I\),
This is the support-domain form of the common-resolvent equation \((\mathrm{basiceq})\) in Section 3.1 of Jenčová–Ruskai, arXiv:0903.2895v4. Its residual calculation is the one in the Appendix, equations \((\mathrm{Mj})\) and \((\mathrm{eq:Schz1})\).
Lemma 21.4.19 writes (??) as a finite sum of non-negative quadratic forms. Hence \(G_i(b_i-S_ix)=0\) for every \(i\). Multiplication by \(S_i\) shows that \(P_i(b_i-S_ix)=0\). Both \(b_i\) and \(S_ix\) lie in the support of \(S_i\), so \(b_i=S_ix\). Multiplication by \(G_i\) now gives (??).
Let \(\rho \) and \(\sigma \) be positive definite on \(\mathcal H_S\otimes \mathbb C^{d_C}\), let \(q=d_C^{-2}\), and put
where \(U_g=\mathbf1_S\otimes W_g\). For \(t{\gt}0\), let
Then
This is the positive-definite, fixed-\(t\) identity in Jenčová–Ruskai, arXiv:0903.2895v4, Appendix, lines 1313–1343. It does not infer zero defect from equality of relative entropies and makes no assertion about singular supports.
Write \(S_g\) for the positive definite matrix representing \(T_g\) under the vectorization \(X\mapsto \operatorname{vec}(X^{\mathsf T})\), and put \(b_g=\operatorname{vec}(B_g^{\mathsf T})\) and \(x=S^{-1}\sum _g b_g\). For the residual \(r_g=b_g-S_gx\), direct expansion gives
Since \(S_g^{1/2}\) is invertible,
Summing proves the identity.
Under the notation of Theorem 21.4.21, set \(\Gamma _t=T^{-1}(A)\). Then, for every \(t{\gt}0\),
This is the second defect family in Jenčová–Ruskai, arXiv:0903.2895v4, §4 and Appendix.
Apply the residual identity of Theorem 21.4.21 with \(a_g=\operatorname{vec}(A_g^{\mathsf T})\) and \(\bar a=\operatorname{vec}(A^{\mathsf T})\). Thus
which is the asserted source-\(A\) identity.
For the positive definite finite-Weyl family, let
and define \(\operatorname{Def}_B(t)\) analogously, with the real parts of the corresponding source-\(B\) pairings. Then both defects are non-negative and continuous for \(t{\gt}0\), and
The integrand is continuous, non-negative, and integrable on \((0,\infty )\). The coefficient of the source-\(B\) defect is exactly \(t\), as prescribed by the integral formula and equality analysis of Jenčová–Ruskai, arXiv:0903.2895v4, §4.
Apply Theorem 21.4.18 to each pair \((A_g,B_g)\) and to \((A,B)\), and interchange the finite sum with the integral. The trace terms cancel because \(B=\sum _gB_g\). The two fixed-resolvent identities write the defects as sums of squared norms, proving nonnegativity. Continuity and integrability follow from the corresponding spectral assertions before taking the finite difference.
Under the hypotheses and notation of the preceding theorem, suppose that
Then \(T_g^{-1}(B_g)=T^{-1}(B)\) for every Weyl index \(g\). In particular this holds for the identity Weyl element \(g=(0,0)\). This is the common-resolvent conclusion in Jenčová–Ruskai, arXiv:0903.2895v4, §4, lines 652–674; its squared-defect input is in Appendix, lines 1313–1343. This fixed-\(t\), positive-definite conclusion does not assert that equality of relative entropies implies the scalar hypothesis in (??).
The hypothesis (??) is a finite sum of squared norms. Each term is non-negative, so every term vanishes. Thus \(S_g^{-1/2}(b_g-S_gx)=0\) for every \(g\). Invertibility of \(S_g^{-1/2}\) gives \(b_g=S_gx\), and hence \(S_g^{-1}b_g=x=S^{-1}\sum _hb_h\).
Under the positive-definite finite-Weyl hypotheses, suppose that \(\sum _gD(A_g\Vert B_g)-D(A\Vert B)=0\). Then \(\operatorname{Def}_B(t)=0\) for every \(t{\gt}0\), and consequently \(T_g^{-1}(B_g)=T^{-1}(B)\) for every \(t{\gt}0\) and every Weyl index \(g\). In particular, equality of relative entropy under the right partial trace implies this conclusion by Theorem 21.4.12. This is the positive-definite conclusion of the equality argument in Jenčová–Ruskai, arXiv:0903.2895v4, §4 and Appendix. It makes no assertion at \(t=0\) or for singular inputs.
The integrand in Theorem 21.4.23 is non-negative and has integral zero, hence it vanishes almost everywhere. Its continuity improves this to vanishing at every \(t{\gt}0\). Since \(\operatorname{Def}_A(t)\geq 0\), \(\operatorname{Def}_B(t)\geq 0\), and \(t/(1+t){\gt}0\), it follows that \(\operatorname{Def}_B(t)=0\). Theorem 21.4.24 now gives the common solution.
Let \(S,T\) be positive semidefinite matrices and let \(x\) be a vector. If \((t\mathbf1+S)^{-1}x=(t\mathbf1+T)^{-1}x\) for every \(t{\gt}0\), then \(\sqrt S\, x=\sqrt T\, x\). More generally, for any fixed matrix \(Q\), if \(Q(t\mathbf1+S)^{-1}x=Q(t\mathbf1+T)^{-1}x\) for every \(t{\gt}0\), then \(Q\sqrt S\, x=Q\sqrt T\, x\).
In particular, for positive definite \(A,B\) and every \(t{\gt}0\), put \(\Delta _{A,B}=A\otimes (B^{-1})^{\mathsf T}\). The source-\(B\) left–right resolvent satisfies
and
These are the positive-square-root specializations of the passage from relative modular resolvents to analytic functions of the relative modular operator in Jenčová–Ruskai, arXiv:0903.2895v4, lines 658–680.
Use the Löwner integral representation of the power \(p=1/2\). Its integrand at \(t{\gt}0\) is \(f_t(S)=t^{-1/2}\mathbf1-t^{1/2}(t\mathbf1+S)^{-1}\). Applied to \(x\), this is \(f_t(S)x=t^{-1/2}x-t^{1/2}(t\mathbf1+S)^{-1}x\). The hypothesis therefore gives \(Q(f_t(S)x)=Q(f_t(T)x)\) for every \(t{\gt}0\). Since \(M\mapsto Q(Mx)\) is a bounded linear map from \(M_{n}(\mathbb {C})\) to \(\mathbb C^n\),
and similarly for \(T\). Integration therefore gives \(Q\sqrt S\, x=Q\sqrt T\, x\).
For the second assertion, factor \(L_A+tR_B=(\Delta _{A,B}+t\mathbf1)R_B\). Applying the inverse to \(B=R_B(\mathbf1)\) gives the shifted relative modular resolvent on \(\mathbf1\). Finally,
as follows from uniqueness of the positive square root.
For matrices \(M\) and \(P\), column-stacking vectorization satisfies
This is the Kronecker vectorization identity \((B\otimes A)\operatorname{vec}(X)=\operatorname{vec}(AXB^{\mathsf T})\) with \(A=P^{\mathsf T}\), \(B=\mathbf1\), and \(X=M^{\mathsf T}\).
Let \(A,B\) be positive semidefinite, let \(t{\gt}0\), set \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\), and let \(P_B\) be the support projection of \(B\). Then
Put \(C=\mathbf1\otimes B^{\mathsf T}\). The support generalized-inverse identity gives
Together with \(B^{\mathsf T}P_B^{\mathsf T}=B^{\mathsf T}\), this factors the left two operators as
The shifted relative-modular matrix is positive definite and hence invertible, so
The Kronecker vectorization identity gives \(C\operatorname{vec}(\mathbf1^{\mathsf T})=\operatorname{vec}(B^{\mathsf T})\), which is the claim.
Let \(A,B\) be positive semidefinite, let \(t{\gt}0\), set
and let \(P_B\) be the support projection of \(B\). Then
Consequently,
This is the one-pair algebraic identification used in the singular equality argument of Jenčová–Ruskai, arXiv:0903.2895v4, lines 783–790. It does not assert that equality of relative entropies gives a common resolvent for a finite family.
Put
The generalized-inverse identities give \(SP=CR\), \(CD=P\), and \(P^2=P\). The preceding lemma shows that \(y=PR^{-1}\operatorname{vec}(\mathbf1^{\mathsf T})\) satisfies \(Sy=\operatorname{vec}(B^{\mathsf T})\), and \(Py=y\). Every vector \(v\) satisfying \(Pv=v\) lies in the range of \(S\), since
In particular, \(y\) lies in the range of \(S\), so the support projection \(P_S\) of \(S\) satisfies \(P_Sy=y\). Therefore
Pairing this vector equality with \(\operatorname{vec}(B^{\mathsf T})\) and taking real parts gives the quadratic identity.
Let \(A\) and \(B\) be positive semidefinite, and write \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\). Then
The support inverse square root is positive semidefinite, and
Uniqueness of the positive square root gives the corresponding Kronecker factorization. Applying the Kronecker vectorization identity to \(\operatorname{vec}(\mathbf1^{\mathsf T})\) gives the stated equation.
Let \(A,B,C,D\) be positive semidefinite matrices. Write \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\) and \(D^+=(D^{-1/2}_{\mathrm{supp}})^2\), and let \(P_B\) be the orthogonal projection onto \((\ker B)^\perp \). If, for every \(t{\gt}0\),
where the inverses act as relative-modular superoperators on matrices, then
This is the square-root specialization of the support functional-calculus passage in Jenčová–Ruskai, arXiv:0903.2895v4, lines 788–793. The projection \(P_B\) is essential: the source gives the common generalized resolvents only after restriction to \((\ker B)^\perp \). Deriving this restricted equality requires the preceding singular equality argument and its kernel hypotheses.
By Lemma 21.4.30,
Setting \(Q=\mathbf1\otimes P_B^{\mathsf T}\) and applying Lemma 21.4.26 gives
Since \(\operatorname{vec}(X^{\mathsf T})=\operatorname{vec}(Y^{\mathsf T})\) implies \(X=Y\), injectivity of vectorization gives the conclusion.
Let \(\rho \) and \(\sigma \) be positive definite and suppose that \(D(\rho \Vert \sigma ) =D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). Put
Then
Hence the raw partial-trace Petz map satisfies \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \). This is the positive-definite case only; no singular-support conclusion is asserted.
Apply the common-resolvent theorem to the identity Weyl summand and the Weyl average. Lemma 21.4.26 converts the result to equality of the two square-root ratios. For \(c=d_C^{-2}\), the uniform scalar in the identity summand cancels through
The Weyl twirl identifies the average with the displayed maximally mixed extensions.
Taking the adjoint product of the ratio equality yields
Multiplication by \(\sqrt\sigma \) on both sides gives the sandwich. The raw Petz recovery identity follows from Theorem 21.4.34.
Let \(\tau \) be positive semidefinite on \(\mathcal H_S\), let \(X\) be a matrix on \(\mathcal H_S\), and let \(\overline\tau =\tau \otimes d_C^{-1}\mathbf1_C\). Then
By Lemma 21.2.43, each outer factor contributes \(\sqrt{d_C}\). The scalar coefficient cancels because \(\sqrt{d_C}\, d_C^{-1}\sqrt{d_C} =d_C^{-1}(\sqrt{d_C})^2=1\). The remaining matrix product is the asserted unital tensor embedding.
Let \(\rho \) and \(\sigma \) be matrices on \(\mathcal H_S\otimes \mathbb C^{d_C}\), with \(\sigma \) positive semidefinite, and set \(\overline\sigma =(\operatorname{tr}_C\sigma )\otimes d_C^{-1}\mathbf1_C\). If the identity summand obeys the support sandwich identity
then the raw support Petz map recovers \(\rho \): \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \). This is the algebraic reduction in [ HJPW04 , Theorem 3, equation (8) ] . It does not derive the support sandwich identity from equality of relative entropies.
Lemma 21.4.33 rewrites the middle three factors as
The support Petz formula therefore identifies the left-hand side of the assumed sandwich identity with \(\mathcal R_\sigma (\operatorname{tr}_C\rho )\).
For every positive semidefinite operator \(\omega _{XY}\),
No invertibility assumption is made on either marginal. This is [ HJPW04 , Equation (4) ] .
The lifted support projections \(P_X\otimes \mathbf1_Y\) and \(\mathbf1_X\otimes P_Y\) both fix \(\omega _{XY}\). Substituting Lemma 21.4.7 into the cross term of the relative entropy and using the two partial-trace adjoint identities gives
The asserted formula follows from \(D(\rho \, \| \sigma )=-S(\rho )-\operatorname{Re}\operatorname{tr}(\rho \log \sigma )\).
Let \(\rho _{ABC}\) be a tripartite density matrix. Then equality holds in strong subadditivity if and only if relative-entropy data processing under the partial trace over \(C\) is saturated for the pair \(\rho _{ABC}\) and \(\rho _A\otimes \rho _{BC}\):
Moreover, \(\operatorname{tr}_C(\rho _A\otimes \rho _{BC})=\rho _A\otimes \rho _B\). This is the product-marginal formulation in [ HJPW04 , Equations (5)–(7) ] .
The two relative entropies are
respectively. Cancelling the common term \(S(\rho _A)\) shows that their equality is precisely
Entrywise, the partial-trace identity is
Let \(\rho _{ABC}\) be a tripartite density matrix and set \(\sigma _{ABC}=(\mathbf1_A/d_A)\otimes \rho _{BC}\). Then equality holds in strong subadditivity if and only if relative-entropy data processing under the partial trace over \(C\) is saturated for this pair:
By Lemma 21.4.3, the reference on the right is \((\mathbf1_A/d_A)\otimes \rho _B\).
This is an equivalent hypothesis-free criterion with a different reference state. The exact product-marginal formulation is Theorem 21.4.36.
The two relative entropies are
respectively. The marginal support lemma places both singular references in the relative-entropy domain. Cancelling the common \(\log d_A\) shows that equality of the two relative entropies is precisely
Let \(\rho \) be a Hermitian operator on \(A\otimes B\otimes C\), and let \(e_A:A'\to A\), \(e_B:B'\to B\), and \(e_C:C'\to C\) be bijections. Define \(\rho '\) by
If \(\rho \) satisfies equality in strong subadditivity, then so does \(\rho '\).
The four operators entering the equality are related by the induced bijections:
These identities follow by changing variables in the three finite sums defining the partial traces. Entropy invariance under reindexing makes the corresponding four entropy terms equal, so the strong-subadditivity equality for \(\rho \) gives the equality for \(\rho '\).
A Hayashi Markov decomposition of a tripartite state \(\rho _{ABC}\) consists of a finite direct-sum decomposition
together with a unitary change of basis on \(B\), a probability vector \((p_j)_j\), and density matrices \(\rho _{A B_j^L}\) and \(\rho _{B_j^R C}\) such that, in the adapted basis, the state becomes
The terminology follows Hayashi’s presentation of quantum Markov structure [ Hay06 ] ; the block decomposition used by the MPDO argument is the structure theorem of [ HJPW04 ] .
For a tripartite density matrix \(\rho _{ABC}\),
holds if and only if \(\rho _{ABC}\) admits a quantum Markov decomposition on the middle subsystem \(B\).
The reverse implication, that a quantum Markov decomposition forces the equality, is proved as Theorem 21.4.41. For the forward implication, the raw product-reference Petz map and its support behavior are given by Theorems 21.2.49 and 21.2.50. The singular-support equality-to-recovery theorem remains separate. After that analytic step, one still needs the family-level Koashi–Imoto decomposition of the recovered states, including the action of the middle-system channel on the common direct-sum factors. This forward implication remains axiomatic. The biconditional combines the two directions.
This equality criterion is cited here as an input for the later MPDO arguments. It is the equality criterion needed by Appendix C of arXiv:1606.00608. The equality conditions are reviewed in [ Hay06 ; Rus02 ] ; the structural block-decomposition formulation used for MPDOs appears in [ HJPW04 ] .
A tripartite density matrix \(\rho _{ABC}\) that admits a quantum Markov decomposition on the middle subsystem \(B\) satisfies
Write the state in the adapted basis as the block-diagonal direct sum \(\bigoplus _j p_j\, \rho _{A B_j^L}\otimes \rho _{B_j^R C}\). The von Neumann entropy of a weighted orthogonal direct sum is \(S(\bigoplus _j p_j\, \omega _j) =-\sum _j p_j\log p_j+\sum _j p_jS(\omega _j)\) by Theorems 21.2.81 and 21.2.80, and the entropy of a tensor product is additive, \(S(\omega \otimes \tau )=S(\omega )+S(\tau )\), by Theorem 21.2.79. Tracing out one tensor factor within each block, the three reduced states factor as block-diagonal direct sums,
with \(\rho _{ABC}\cong \bigoplus _j p_j\, \rho _{AB_j^L}\otimes \rho _{B_j^R C}\) itself. Writing \(H=-\sum _j p_j\log p_j\), the four entropies expand as
Both sides of the claimed identity equal
so \(S(\rho _{ABC})+S(\rho _B)=S(\rho _{AB})+S(\rho _{BC})\). The basis change on \(B\) and the direct-sum reindexing leave every entropy term unchanged.
21.5 Mutual information
The quantum mutual information of a bipartite state \(\rho _{AB}\) is
For any bipartite density matrix \(\rho _{AB}\), \(I(A{:}B) \ge 0\).
Apply strong subadditivity (Theorem 21.4.4) with trivial \(B\) (one-dimensional middle system). The SSA inequality \(S(\rho _{ABC}) + S(\rho _B) \le S(\rho _{AB}) + S(\rho _{BC})\) reduces to subadditivity \(S(\rho _{AC}) \le S(\rho _A) + S(\rho _C)\), which gives \(I(A{:}C) \ge 0\).
21.6 Entropy formulations
This section states the entropy formulations used later in the development: von Neumann entropy, strong subadditivity, quantum Markov decomposition, and mutual information. These statements are cited from the standard entropy literature and supply the entropy-theoretic input for the later MPDO arguments.
This formulation has the same value \(S(\rho ) = -\sum _i \lambda _i \log \lambda _i\) as in Definition 21.2.1.
For any tripartite density matrix \(\rho _{ABC}\) on \(A \otimes B \otimes C\),
This formulation introduces no new axiom: it is the same strong-subadditivity statement as Theorem 21.4.4, which is proved there from Lieb concavity along the relative-entropy route [ LR73 ] .
This is the entropy formulation of the Hayashi Markov decomposition from Definition 21.4.39.
For any tripartite density matrix \(\rho _{ABC}\), equality in strong subadditivity holds if and only if \(\rho _{ABC}\) admits a quantum Markov decomposition on the middle subsystem \(B\).
This formulation introduces no new axiom: it is the same equality criterion as Theorem 21.4.40.
This formulation has the same value \(I(A{:}B) = S(\rho _A) + S(\rho _B) - S(\rho _{AB})\) as in Definition 21.5.1.
For a tripartite density matrix \(\rho _{ABC}\) with \(\dim B = 1\), one has \(S(\rho _{ABC}) \le S(\rho _{AB}) + S(\rho _{BC})\).
21.7 Mutual information: monotonicity and area-law bound
The two inequalities below are the downstream MPDO-facing consequences of the entropy inequalities in this chapter. The monotonicity inequality is the strong-subadditivity content underlying the MPDO mutual-information monotonicity \(I_L \le I_{L+1}\) (arXiv:1606.00608, Proposition C.1); the elementary area-law bound is the single-site entropy bound underlying the MPDO area-law bound \(I_L \le 4\log D\) (arXiv:1606.00608, cited from the Wolf area-law bound). Both inequalities follow directly from strong subadditivity (Theorem 21.4.4) and the single-system \(S(\rho ) \le \log D\) bound; neither introduces a new axiom.
For any PSD Hermitian matrix \(\rho \) with \(\operatorname{tr}(\rho ) = 1\) on an arbitrary finite index set, \(S(\rho ) \ge 0\). This is the analog of Theorem 21.2.4 for arbitrary finite index sets: it does not require the index set to be \(\{ 0,\ldots ,D{-}1\} \), so it applies to bipartite density matrices on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B}\).
The eigenvalues of a PSD matrix are non-negative, and eigenvalues of a trace-\(1\) Hermitian matrix with non-negative eigenvalues are bounded above by \(1\) (each single eigenvalue is at most the total sum). Apply \(x\log x \le 0\) on \([0, 1]\) pointwise and sum.
For any tripartite density matrix \(\rho _{ABC}\) on \(A \otimes B \otimes C\),
where the left-hand side is the bipartite mutual information of the reduced state \(\rho _{AB} = \operatorname{tr}_C(\rho _{ABC})\) and the right-hand side is evaluated by expanding \(I(A{:}BC) = S(\rho _A) + S(\rho _{BC}) - S(\rho _{ABC})\).
Expanding both sides in entropy form and cancelling the common \(S(\rho _A)\) term, the inequality reduces to strong subadditivity \(S(\rho _{ABC}) + S(\rho _B) \le S(\rho _{AB}) + S(\rho _{BC})\). The bipartite \(B\)-reduced state of \(\rho _{AB}\) matches the tripartite \(B\)-reduced state \(\operatorname{tr}_{AC}(\rho _{ABC})\), ensuring the \(S(\rho _B)\) term is the one supplied by Theorem 21.6.2.
For a bipartite density matrix \(\rho _{AB}\) on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B}\) with \(d_A, d_B \ge 1\) whose single-system reduced states are obtained by partial trace, \(I(A{:}B) \le \log d_A + \log d_B\).
21.8 Trivial-factor corollaries
The partial traces and von Neumann entropy are introduced in Chapter 21. This supplement collects elementary trace-preservation identities and the dimension-1 entropy bound that follow directly from those definitions. They are listed as separate results because each proof requires only unfolding definitions and re-indexing finite sums.
For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_A(\rho _{ABC}))\).
Unfolding the definitions,
For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_C(\rho _{ABC}))\).
Unfolding the definitions,
For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_{AC}(\rho _{ABC}))\).
Unfolding the definitions and interchanging the outer sums,
If \(\rho \in M_{1}(\mathbb {C})\) is Hermitian and \(\operatorname{tr}(\rho )=1\), then \(S(\rho )=0\).
The unique eigenvalue equals the trace, hence equals \(1\), and \(-1\cdot \log 1=0\):