Quantum Information and Channels: A formalization blueprint

13 Quantum Entropy

This chapter develops trace distance, von Neumann entropy, and their basic properties for finite-dimensional quantum systems.

13.1 Trace norm

Definition 13.1.1 Schatten one-norm
#

For a matrix \(A\in M_{D}(\mathbb {C})\), the singular values \(s_0(A),s_1(A),\ldots \) form a finitely supported family: only finitely many are nonzero. The Schatten one-norm of \(A\) is their sum, that is, the sum of the finitely many nonzero singular values:

\begin{align} \lVert A\rVert _1 & =\sum _{i\, :\, s_i(A)\neq 0}s_i(A). \label{eq:entropy_schatten_one_norm} \end{align}

This is the \(p=1\) case of the Schatten \(p\)-norm [ Wol12 , Chapter 8, Section 8.1 ] . The equivalent closed formula \(\lVert A\rVert _1=\sum _{i=0}^{D-1}s_i(A)\), summing over all \(D\) singular values including trailing zeros, is part of Theorem 13.1.4.

Definition 13.1.2 Trace norm
#

The trace norm of \(A\in M_{D}(\mathbb {C})\) is its Schatten one-norm:

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\lVert A\rVert _1. \label{eq:entropy_trace_norm} \end{align}

Wolf records the equivalent formula \(\lVert A\rVert _1=\operatorname{tr}\lvert A\rvert \) [ Wol12 , Chapter 8, Section 8.1 ] ; this is Theorem 13.1.5.

The Schatten one-norm and trace norm are the sums over the finite support of the singular-value sequence. Equivalently, the trace norm is the sum over the indices below the rank of the represented linear map.

Proof

The singular-value sequence satisfies \(s_i(A)=0\) for all \(i\geq \operatorname{rank}(A)\), so summing over the support and summing over \(\{ 0,\ldots ,\operatorname{rank}(A)-1\} \) give the same value.

For every \(A\in M_{D}(\mathbb {C})\),

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\sum _{i=0}^{D-1}s_i(A), \label{eq:entropy_trace_norm_fin}\\ \lVert A\rVert _{\operatorname{tr}} & \geq 0. \notag \end{align}

Also, \(\lVert A\rVert _{\operatorname{tr}}=0\) if and only if \(A=0\). Equivalently, \(\lVert A\rVert _{\operatorname{tr}}{\gt}0\) if and only if \(A\neq 0\).

Proof

The singular values satisfy \(s_i(A)\geq 0\), so \(\lVert A\rVert _{\operatorname{tr}}=0 \iff \forall i,\, s_i(A)=0 \iff A=0\).

Theorem 13.1.5 Trace norm as trace of the absolute value

For every \(A\in M_{D}(\mathbb {C})\), with \(\lvert A\rvert =\sqrt{A^\dagger A}\) the positive-semidefinite square root of \(A^\dagger A\),

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\operatorname{Re}(\operatorname{tr}\lvert A\rvert ). \label{eq:entropy_trace_norm_abs} \end{align}

This is the formula \(\lVert A\rVert _1=\operatorname{tr}[\lvert A\rvert ]\) of [ Wol12 , Chapter 8, Section 8.1 ] ; the trace of the positive-semidefinite matrix \(\lvert A\rvert \) is real.

Proof

The map \(T^\dagger T\), for \(T\) the linear map on \(\mathbb {C}^D\) represented by \(A\), is represented by \(A^\dagger A\), so the singular values of \(A\) are \(s_i(A)=\sqrt{\lambda _i}\) with \(\lambda _0,\ldots ,\lambda _{D-1}\) the eigenvalues of \(A^\dagger A\). Diagonalizing \(A^\dagger A=U\operatorname{diag}(\lambda _i)U^\dagger \) gives \(\lvert A\rvert =U\operatorname{diag}(\sqrt{\lambda _i})U^\dagger \), hence

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\sum _{i=0}^{D-1}s_i(A) =\sum _{i=0}^{D-1}\sqrt{\lambda _i} =\operatorname{tr}\lvert A\rvert =\operatorname{Re}(\operatorname{tr}\lvert A\rvert ), \notag \end{align}

where the last step uses that \(\lvert A\rvert \) is positive semidefinite, so \(\operatorname{tr}\lvert A\rvert \geq 0\) is real.

Theorem 13.1.6 Trace-norm homogeneity

For every \(c\in \mathbb {C}\) and \(A\in M_{D}(\mathbb {C})\), \(\lVert cA\rVert _{\operatorname{tr}} =\lvert c\rvert \, \lVert A\rVert _{\operatorname{tr}}\). This is the homogeneity axiom for matrix norms [ Wol12 , Chapter 8, Section 8.1 ] .

Proof

From \((cA)^\dagger (cA)=\lvert c\rvert ^2\, A^\dagger A\) and uniqueness of the positive-semidefinite square root, \(\lvert cA\rvert =\lvert c\rvert \, \lvert A\rvert \). Linearity of the trace gives \(\operatorname{tr}\lvert cA\rvert =\lvert c\rvert \, \operatorname{tr}\lvert A\rvert \), and the claim follows from Theorem 13.1.5.

For all unitaries \(U,V\in M_{D}(\mathbb {C})\) and every \(A\in M_{D}(\mathbb {C})\),

\begin{align} \lVert UA\rVert _{\operatorname{tr}} & =\lVert AV\rVert _{\operatorname{tr}} =\lVert UAV\rVert _{\operatorname{tr}} =\lVert A\rVert _{\operatorname{tr}}. \notag \end{align}

The trace norm is thus unitarily invariant [ Wol12 , Chapter 8, Section 8.1 ] .

Proof

From \((UA)^\dagger (UA)=A^\dagger U^\dagger UA=A^\dagger A\) the absolute values agree: \(\lvert UA\rvert =\lvert A\rvert \). For the right factor, \((AV)^\dagger (AV)=V^\dagger (A^\dagger A)V\), and since \(V^\dagger \lvert A\rvert V\) is positive semidefinite with

\begin{align} (V^\dagger \lvert A\rvert V)^2 & =V^\dagger \lvert A\rvert \, VV^\dagger \, \lvert A\rvert V =V^\dagger (A^\dagger A)V, \notag \end{align}

uniqueness of the positive-semidefinite square root gives \(\lvert AV\rvert =V^\dagger \lvert A\rvert V\). Cyclicity of the trace then yields \(\operatorname{tr}\lvert AV\rvert =\operatorname{tr}(\lvert A\rvert \, VV^\dagger ) =\operatorname{tr}\lvert A\rvert \). For the two-sided case, \(\lVert UAV\rVert _{\operatorname{tr}} =\lVert UA\rVert _{\operatorname{tr}} =\lVert A\rVert _{\operatorname{tr}}\) by applying the right and left cases in turn.

Lemma 13.1.8 Trace norm from the spectrum of \(A^\dagger A\)

For every \(A\in M_{D}(\mathbb {C})\), the trace norm is the sum of the square roots of the eigenvalues of the positive operator \(A^\dagger A\):

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\sum _{i=0}^{D-1}\sqrt{\lambda _i(A^\dagger A)}. \notag \end{align}
Proof

Since \(s_i(A)=\sqrt{\lambda _i(A^\dagger A)}\) by the definition of singular value, the finite singular-value expansion (3) gives

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\sum _{i=0}^{D-1}s_i(A) =\sum _{i=0}^{D-1}\sqrt{\lambda _i(A^\dagger A)}. \notag \end{align}
Lemma 13.1.9 Trace norm from the matrix eigenvalues of \(A^\dagger A\)

For every \(A\in M_{D}(\mathbb {C})\), with \(\lambda _0,\ldots ,\lambda _{D-1}\) the eigenvalues of the matrix \(A^\dagger A\), \(\lVert A\rVert _{\operatorname{tr}} =\sum _{i=0}^{D-1}\sqrt{\lambda _i}\).

Proof

Diagonalizing \(A^\dagger A=V\operatorname{diag}(\lambda _i)V^\dagger \) gives \(\lvert A\rvert =V\operatorname{diag}(\sqrt{\lambda _i})V^\dagger \), so \(\operatorname{tr}\lvert A\rvert =\sum _{i=0}^{D-1}\sqrt{\lambda _i}\), and the claim follows from Theorem 13.1.5.

Theorem 13.1.10 Trace norm and Hilbert–Schmidt norm comparison

For every \(A\in M_{D}(\mathbb {C})\),

\begin{align} \lVert A\rVert _2 & \leq \lVert A\rVert _1 \leq \sqrt D\, \lVert A\rVert _2. \label{eq:trace_norm_frobenius_comparison} \end{align}

The first inequality is the \(p=1\), \(p'=2\) case of [ Wol12 , Eq. (8.1) ] ; the second is the corresponding case of [ Wol12 , Eq. (8.7) ] . Both are used in the proof of trace-norm convergence toward asymptotic states.

Proof

Write \(s_0,\ldots ,s_{D-1}\geq 0\) for the singular values of \(A\). Then \(\lVert A\rVert _2^2=\sum _i s_i^2\) and \(\lVert A\rVert _1=\sum _i s_i\). The first inequality follows by expanding \((\sum _i s_i)^2\) and using nonnegativity. The second is the Cauchy–Schwarz bound \((\sum _i s_i)^2\leq D\sum _i s_i^2\).

For every \(A\in M_{D}(\mathbb {C})\),

\begin{align} \lVert A\rVert _{\operatorname{tr}} & =\max \left\{ \bigl|\operatorname{tr}[A^\dagger U]\bigr| \middle | UU^\dagger =\mathbb {1}\right\} , \label{eq:entropy_trace_norm_variational} \end{align}

and the maximum is attained: some unitary \(U\) satisfies \(\operatorname{tr}[A^\dagger U]=\lVert A\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Eq. (8.11) ] .

Proof

For the upper bound, let \(v_0,\ldots ,v_{D-1}\) be an orthonormal eigenbasis of \(A^\dagger A\) with eigenvalues \(\lambda _i\), and let \(U\) be unitary. Expanding the trace in this basis and applying the Cauchy–Schwarz inequality,

\begin{align} \bigl|\operatorname{tr}[A^\dagger U]\bigr| & =\Bigl|\sum _i\langle Av_i,\, Uv_i\rangle \Bigr| \leq \sum _i\lVert Av_i\rVert \, \lVert Uv_i\rVert =\sum _i\sqrt{\lambda _i} =\lVert A\rVert _{\operatorname{tr}}, \notag \end{align}

since \(\lVert Av_i\rVert ^2=\langle v_i,\, A^\dagger A\, v_i\rangle =\lambda _i\) and \(\lVert Uv_i\rVert =1\).

For attainment, the vectors \(w_i=Av_i/\sqrt{\lambda _i}\), taken over the indices with \(\lambda _i\neq 0\), satisfy

\begin{align} \langle w_i,w_j\rangle & =\frac{\langle v_i,\, A^\dagger A\, v_j\rangle }{\sqrt{\lambda _i}\sqrt{\lambda _j}} =\delta _{ij}, \notag \end{align}

so they form an orthonormal family. Extend it to an orthonormal basis \((w_i)_i\) of \(\mathbb {C}^D\) and let \(U\) be the unitary with \(Uv_i=w_i\) for all \(i\). Then

\begin{align} \operatorname{tr}[A^\dagger U] & =\sum _i\langle Av_i,\, w_i\rangle =\sum _{\lambda _i\neq 0}\frac{\lVert Av_i\rVert ^2}{\sqrt{\lambda _i}} =\sum _i\sqrt{\lambda _i} =\lVert A\rVert _{\operatorname{tr}}, \notag \end{align}

where the indices with \(\lambda _i=0\) contribute \(0\) to both sides because \(\lVert Av_i\rVert ^2=\lambda _i=0\) forces \(Av_i=0\).

Theorem 13.1.12 Trace-norm triangle inequality

For all \(A,B\in M_{D}(\mathbb {C})\), \(\lVert A+B\rVert _{\operatorname{tr}} \leq \lVert A\rVert _{\operatorname{tr}}+\lVert B\rVert _{\operatorname{tr}}\). Together with homogeneity (Theorem 13.1.6) and definiteness (Theorem 13.1.4), this completes the norm axioms of [ Wol12 , Chapter 8, Section 8.1 ] for the trace norm.

Proof

Choose by Theorem 13.1.11 a unitary \(U\) with \(\operatorname{tr}[(A+B)^\dagger U]=\lVert A+B\rVert _{\operatorname{tr}}\). Splitting the trace and applying the upper-bound half of the same theorem to \(A\) and to \(B\) separately,

\begin{align} \lVert A+B\rVert _{\operatorname{tr}} & =\bigl|\operatorname{tr}[A^\dagger U] +\operatorname{tr}[B^\dagger U]\bigr| \leq \bigl|\operatorname{tr}[A^\dagger U]\bigr| +\bigl|\operatorname{tr}[B^\dagger U]\bigr| \leq \lVert A\rVert _{\operatorname{tr}}+\lVert B\rVert _{\operatorname{tr}}. \notag \end{align}
Lemma 13.1.13 Jordan trace-norm formula

Let \(H\in M_{D}(\mathbb {C})\) be Hermitian with Jordan decomposition \(H=H^+-H^-\) into orthogonal positive parts, \(H^\pm \geq 0\) and \(H^+H^-=0\). Then \(\lVert H\rVert _{\operatorname{tr}}=\operatorname{tr}[H^+]+\operatorname{tr}[H^-]\).

Proof

The sum \(H^++H^-\) is positive semidefinite, and by orthogonality its square is

\begin{align} (H^++H^-)^2 & =(H^+)^2+(H^-)^2 =(H^+-H^-)^2 =H^\dagger H, \notag \end{align}

so \(H^++H^-=\sqrt{H^\dagger H}=\lvert H\rvert \). Hence \(\lVert H\rVert _{\operatorname{tr}} =\operatorname{tr}\lvert H\rvert =\operatorname{tr}[H^+]+\operatorname{tr}[H^-]\) by Theorem 13.1.5.

Let \(A\in M_{D}(\mathbb {C})\) be Hermitian with positive part \(A^+\). There is a matrix \(\Pi \) with \(0\leq \Pi \leq \mathbb {1}\), \(\Pi ^2=\Pi \), and \(\Pi A=A^+\), namely the orthogonal projection onto the support space of \(A^+\).

Proof

Diagonalize \(A=U\operatorname{diag}(\lambda _1,\ldots ,\lambda _D)\, U^\dagger \) and set

\begin{align} \Pi & =U\operatorname{diag}(\chi (\lambda _1),\ldots ,\chi (\lambda _D))\, U^\dagger , \notag \end{align}

where \(\chi \) is the indicator function of \((0,\infty )\). Each claim is read off eigenvalue-wise: \(0\leq \chi \leq 1\) gives \(0\leq \Pi \leq \mathbb {1}\), \(\chi ^2=\chi \) gives \(\Pi ^2=\Pi \), and \(\chi (\lambda ) \lambda =\max (\lambda ,0)\) gives \(\Pi A=A^+\).

Lemma 13.1.15 Positive-semidefinite trace pairing bounds

Let \(X,C\in M_{D}(\mathbb {C})\) with \(X\geq 0\). If \(C\geq 0\), then \(\operatorname{tr}[CX]\geq 0\); if \(C\leq \mathbb {1}\), then \(\operatorname{tr}[CX]\leq \operatorname{tr}[X]\).

Proof

The first bound is Lemma 9.5.1. For the second, \(\operatorname{tr}[X]-\operatorname{tr}[CX]=\operatorname{tr}[(\mathbb {1}-C)X]\geq 0\) by the same lemma, since \(\mathbb {1}-C\geq 0\).

Lemma 13.1.16 Positive maps preserve hermiticity

If \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) is a positive linear map and \(H\in M_{D}(\mathbb {C})\) is Hermitian, then \(T(H)\) is Hermitian.

Proof

Write \(H=H^+-H^-\) with \(H^\pm \geq 0\). Then \(T(H)=T(H^+)-T(H^-)\) is a difference of positive semidefinite matrices, hence Hermitian.

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a positive linear map and let \(H\in M_{D}(\mathbb {C})\) be Hermitian. Then \(\operatorname{tr}[(TH)^+]\leq \operatorname{tr}[T(H^+)]\).

Proof

Let \(\Pi _+\) be the support projection of \((TH)^+\) from Lemma 13.1.14, so that \(\Pi _+ T(H) = (T(H))^+\) and \(0\leq \Pi _+\leq \mathbb {1}\). The estimate follows from the chain

\begin{align} \operatorname{tr}[(TH)^+] & = \operatorname{tr}[\Pi _+ T(H^+-H^-)] \notag \\ & = \operatorname{tr}[\Pi _+ T(H^+)] - \operatorname{tr}[\Pi _+ T(H^-)] \notag \\ & \leq \operatorname{tr}[\Pi _+ T(H^+)] \notag \\ & \leq \operatorname{tr}[T(H^+)] , \end{align}

using Lemmas 13.1.15 and 13.1.16.

Lemma 13.1.18 Positive-part trace estimate

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map and let \(H\in M_{D}(\mathbb {C})\) be Hermitian, with Jordan decompositions \(H=P_+-P_-\) and \(T(H)=Q_+-Q_-\). Then \(\operatorname{tr}[Q_+]\leq \operatorname{tr}[P_+]\).

Proof

Lemma 13.1.17 gives \(\operatorname{tr}[(TH)^+]\leq \operatorname{tr}[T(H^+)]\). The result follows because trace preservation gives \(\operatorname{tr}[T(H^+)]=\operatorname{tr}[H^+]\).

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map. Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\lVert T(H)\rVert _{\operatorname{tr}}\leq \lVert H\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Theorem 8.16 ] .

Proof

Write \(H=P_+-P_-\) and \(T(H)=Q_+-Q_-\) for the Jordan decompositions. Applying Lemma 13.1.18 to \(H\) and to \(-H\) gives \(\operatorname{tr}[Q_+]\leq \operatorname{tr}[P_+]\) and \(\operatorname{tr}[Q_-]\leq \operatorname{tr}[P_-]\), so by the Jordan trace-norm formula (Lemma 13.1.13),

\begin{align} \lVert T(H)\rVert _{\operatorname{tr}} & =\operatorname{tr}[Q_+]+\operatorname{tr}[Q_-] \leq \operatorname{tr}[P_+]+\operatorname{tr}[P_-] =\lVert H\rVert _{\operatorname{tr}}. \notag \end{align}

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving positive linear map. Then for all density matrices \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\), \(\lVert T(\rho _1)-T(\rho _2)\rVert _{\operatorname{tr}} \leq \lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}\). See [ Wol12 , Chapter 8, Eq. (8.80) ] .

Proof

By linearity \(T(\rho _1)-T(\rho _2)=T(\rho _1-\rho _2)\), and \(\rho _1-\rho _2\) is Hermitian as a difference of positive semidefinite matrices, so Theorem 13.1.19 applies.

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a complex-linear map. Then

\begin{align} \sup _{\substack {\rho _1,\rho _2\in \mathcal{D}_D\\ \rho _1\neq \rho _2}} \frac{\lVert T(\rho _1)-T(\rho _2)\rVert _{\operatorname{tr}}}{\lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}} =\frac12\sup _{\substack {\lVert \psi \rVert =\lVert \phi \rVert =1\\ \langle \psi ,\phi \rangle =0}} \left\lVert T\left(|\psi \rangle \! \langle \psi | -|\phi \rangle \! \langle \phi |\right)\right\rVert _{\operatorname{tr}}. \label{eq:trace_norm_contraction_coefficient} \end{align}

The left supremum is over distinct density matrices in \(M_{D}(\mathbb {C})\), and the right supremum is over orthogonal unit vectors in \(\mathbb {C}^D\). No positivity or trace-preservation assumption is imposed on \(T\). This is [ Wol12 , Chapter 8, Lemma 8.3, Eq. (8.81) ] .

Proof

An orthogonal pair of pure states gives a pair of distinct density matrices at trace-norm distance \(2\), which proves the lower bound. For the reverse inequality, set \(H=\rho _1-\rho _2\) and write its Jordan decomposition as \(H=H^+-H^-\). Since \(H\) is nonzero and traceless, \(\operatorname{tr}[H^+]=\operatorname{tr}[H^-]=t{\gt}0\). Thus \(P=H^+/t\) and \(Q=H^-/t\) are density matrices with orthogonal supports, and homogeneity reduces the quotient for \(H\) to \(\frac12\lVert T(P-Q)\rVert _{\operatorname{tr}}\).

Take spectral decompositions \(P=\sum _i\lambda _i|\psi _i\rangle \! \langle \psi _i|\) and \(Q=\sum _j\mu _j|\phi _j\rangle \! \langle \phi _j|\). Orthogonality of the supports gives \(\psi _i\perp \phi _j\) whenever \(\lambda _i\mu _j\neq 0\), while \(\lambda _i,\mu _j\geq 0\) and \(\sum _i\lambda _i=\sum _j\mu _j=1\). The product-weight expansion

\begin{align} P-Q=\sum _{i,j}\lambda _i\mu _j \left(|\psi _i\rangle \! \langle \psi _i|-|\phi _j\rangle \! \langle \phi _j|\right) \notag \end{align}

is therefore a convex combination of differences of orthogonal pure states. The triangle inequality and homogeneity bound its image under \(T\) by the right-hand supremum. Taking the supremum over \(\rho _1\neq \rho _2\) proves the equality.

Definition 13.1.22 Trace-norm distance to the asymptotic image

For \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\), let \(T_\phi \) be its peripheral spectral projection and define

\begin{align} \Delta _T(\rho ):=\lVert \rho -T_\phi (\rho )\rVert _{\operatorname{tr}}. \end{align}

This is Wolf’s notation in [ Wol12 , Chapter 8, Proposition “Convergence towards asymptotic states”, Eq. (8.112) ] .

Let \(T_\varphi :=T\circ T_\phi =T_\phi \circ T\) be Wolf’s asymptotic dynamics. For every matrix \(\rho \), every \(n\geq 0\), and every \(n{\gt}0\) in the second equality,

\begin{align} \Delta _T(T^n(\rho )) & =\lVert T^n(\rho -T_\phi (\rho ))\rVert _{\operatorname{tr}}\\ & =\lVert (T^n-T_\varphi ^n)(\rho -T_\phi (\rho ))\rVert _{\operatorname{tr}}. \label{eq:trace_norm_asymptotic_distance_numerator} \end{align}

These are the numerator identities used in [ Wol12 , Chapter 8, Eq. (8.114) ] . They follow from \(T_\phi T=TT_\phi =T_\varphi =T_\varphi T_\phi \); the positive-iterate condition records that the zeroth power of \(T_\varphi \) is the identity.

Proof

The commutation \(T_\phi T=TT_\phi \) iterates to \(T_\phi T^n=T^nT_\phi \), which proves the first equality by linearity. Idempotence of \(T_\phi \) gives \(T_\varphi (\rho -T_\phi (\rho ))=0\); every positive power of \(T_\varphi \) therefore vanishes on the same remainder, proving the second equality.

Lemma 13.1.24 Pointwise form of the contraction-coefficient bound

Let \(S:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be complex-linear and let \(\rho _1\neq \rho _2\) be density matrices. Then

\begin{align} \frac{\lVert S(\rho _1)-S(\rho _2)\rVert _{\operatorname{tr}}}{\lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}} \leq \frac12 \sup _{\substack {\lVert \psi \rVert =\lVert \phi \rVert =1\\ \psi \perp \phi }} \left\lVert S\left(|\psi \rangle \! \langle \psi | -|\phi \rangle \! \langle \phi |\right)\right\rVert _{\operatorname{tr}}. \end{align}

This is the pointwise inequality from Wolf Lemma 8.3 used in Equation (8.115).

Proof

The proof is the upper-bound half of Lemma 13.1.21: normalize the positive and negative parts of \(\rho _1-\rho _2\) to orthogonally supported density matrices and expand both in eigenprojectors.

Lemma 13.1.25 Hilbert–Schmidt estimates for the asymptotic bound

For \(A\in M_{D}(\mathbb {C})\) and a complex-linear map \(S:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\),

\begin{align} \lVert A\rVert _{\operatorname{tr}}& \leq \sqrt D\, \lVert A\rVert _2,\\ \lVert S(A)\rVert _2& \leq \lVert S\rVert _{2\to 2}\lVert A\rVert _2. \end{align}

If \(\psi \) and \(\phi \) are orthogonal unit vectors, then \(\lVert |\psi \rangle \! \langle \psi |-|\phi \rangle \! \langle \phi |\rVert _2=\sqrt2\). Consequently the orthogonal-pure-state supremum for \(S\) is at most \(\sqrt D\, \lVert S\rVert _{2\to 2}\sqrt2\). These are precisely the estimates in Wolf Equations (8.115)–(8.116).

Proof

The first inequality is finite Cauchy–Schwarz applied to the singular values. The second is the defining application bound for the operator norm after Frobenius vectorization. Orthogonality makes the two rank-one projectors multiply to zero, so the squared Hilbert–Schmidt norm of their difference is the sum of their traces, namely \(2\).

Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be positive and trace preserving, let \(\rho \) be a density operator, and let \(n{\gt}0\). Then

\begin{align} \Delta _T(T^n(\rho )) \leq \sqrt{\frac D2}\, \left\lVert \widehat T^{\, n}-\widehat T_\varphi ^{\, n}\right\rVert _\infty \, \Delta _T(\rho ). \tag {8.112} \end{align}

Here the displayed transfer-matrix norm equals \(\lVert T^n-T_\varphi ^n\rVert _{2\to 2}\), the Hilbert–Schmidt operator norm of the corresponding superoperator. The explicit \(n{\gt}0\) records the paper’s positive-natural convention: at \(n=0\), both superoperator powers are the identity, so the displayed right-hand side vanishes. This is the upper assertion of  [ Wol12 , Chapter 8, Proposition “Convergence towards asymptotic states”, Eqs. (8.112), (8.114)–(8.116) ] .

Proof

Apply the numerator identity to \(S=T^n-T_\varphi ^n\). The peripheral projection is again positive and trace preserving, so Lemma 8.3 applies to the distinct density matrices \(\rho \) and \(T_\phi (\rho )\). Then use \(\lVert \cdot \rVert _{\operatorname{tr}}\leq \sqrt D\lVert \cdot \rVert _2\) and the \(\sqrt2\) Hilbert–Schmidt distance of orthogonal pure states. The constants simplify as \(\frac12\sqrt D\sqrt2=\sqrt{D/2}\); if \(\rho =T_\phi (\rho )\), both sides vanish directly. Finally, the positive-power identity identifies the transfer-matrix difference with the transfer matrix of \(T^n-T_\varphi ^n\), whose largest-singular-value norm equals the Hilbert–Schmidt operator norm \(\lVert T^n-T_\varphi ^n\rVert _{2\to 2}\).

Lemma 13.1.27 Scaled positive-part trace estimate

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a positive linear map with \(\operatorname{tr}[T(\rho )] = c\, \operatorname{tr}[\rho ]\) for some nonnegative real \(c\) and all \(\rho \in M_{D}(\mathbb {C})\). Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\operatorname{tr}[(TH)^+]\leq c\cdot \operatorname{tr}[H^+]\). Generalizes Lemma 13.1.18 from the trace-preserving case \(c=1\).

Proof

Lemma 13.1.17 gives \(\operatorname{tr}[(TH)^+]\leq \operatorname{tr}[T(H^+)]\). The trace-scaling hypothesis then yields \(\operatorname{tr}[T(H^+)] = c\, \operatorname{tr}[H^+]\), giving the claimed bound.

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a positive linear map with \(\operatorname{tr}[T(\rho )] = c\, \operatorname{tr}[\rho ]\) for some nonnegative real \(c\) and all \(\rho \in M_{D}(\mathbb {C})\). Then for all Hermitian \(H\in M_{D}(\mathbb {C})\), \(\lVert T(H)\rVert _{\operatorname{tr}}\leq c\cdot \lVert H\rVert _{\operatorname{tr}}\). Generalizes Theorem 13.1.19 from the trace-preserving case \(c=1\).

Proof

Same argument as Theorem 13.1.19: apply Lemma 13.1.27 to \(H\) and \(-H\), then sum using the Jordan trace-norm formula.

Let \(T,T':M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be trace-preserving and Hermiticity-preserving linear maps with \(T'(X)=\operatorname{tr}[X]\, Y\) for some fixed \(Y\in M_{D'}(\mathbb {C})\). If \(T-\varepsilon T'\) is positive for some \(\varepsilon \geq 0\), then for all density operators \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\),

\begin{align} \lVert T(\rho _1)-T(\rho _2)\rVert _{\operatorname{tr}} \leq (1-\varepsilon )\, \lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}. \label{eq:quantum_doeblin_contraction} \end{align}

See [ Wol12 , Chapter 8, Theorem 8.17 ] .

Proof

The map \(S:=T-\varepsilon T'\) is positive by hypothesis. Since \(T'(\rho _1-\rho _2)=\operatorname{tr}[\rho _1-\rho _2]\, Y=0\) (the densities have unit trace), \(T(\rho _1)-T(\rho _2)=S(\rho _1-\rho _2)\). The trace of \(S(\rho )\) is \((\operatorname{tr}[\rho ])-\varepsilon \operatorname{tr}[\rho ]=(1-\varepsilon ) \operatorname{tr}[\rho ]\), so \(S\) satisfies the scaled-trace hypothesis of Theorem 13.1.28 with \(c=1-\varepsilon \). Positivity of \(S\) on \(\rho _1\) forces \(0\leq 1-\varepsilon \) (by PSD trace nonnegativity), so the scaled-trace theorem applies and yields the claimed bound.

Definition 13.1.30 Trace-prepare map

For any matrix \(Y\in M_{D'}(\mathbb {C})\), define \(T'_Y:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) by

\begin{align} T’_Y(X) = \operatorname{tr}[X]\, Y. \label{eq:trace_prepare_map} \end{align}
Lemma 13.1.31 Choi matrix of the trace-prepare map

For arbitrary \(Y\in M_{D'}(\mathbb {C})\), the rectangular (output-factor-first) Choi matrix of \(T'_Y\) is

\begin{align} \tau (T’_Y) = \frac{1}{D}\, (Y\otimes \mathbb {1}). \label{eq:choi_trace_prepare} \end{align}

See [ Wol12 , Chapter 8, Eq. (8.86) ] .

Proof

The normalized omega slice is \(D^{-1}E_{ij}\), where \(E_{ij}\) is the matrix unit. Since \(\operatorname{tr}(E_{ij})=\delta _{ij}\), the trace of the slice is \(D^{-1}\delta _{ij}\), and \(T'_Y\) multiplies this scalar by \(Y\). On the other hand, \((Y\otimes \mathbb {1})_{(a,i),(b,j)} = Y_{ab}\, \delta _{ij}\), so both sides equal \(D^{-1}Y_{ab}\, \delta _{ij}\).

Let \(T:M_{D}(\mathbb {C})\to M_{D'}(\mathbb {C})\) be a trace-preserving, Hermiticity-preserving linear map and \(Y\in M_{D'}(\mathbb {C})\) Hermitian with \(\operatorname{tr}[Y]=1\). For \(\varepsilon \geq 0\), if the rectangular Choi matrix \(\tau (T)\) satisfies

\begin{align} \tau (T) \ge \frac{\varepsilon }{D}\, (Y\otimes \mathbb {1}), \label{eq:choi_domination} \end{align}

then for all density operators \(\rho _1,\rho _2\in M_{D}(\mathbb {C})\),

\begin{align} \lVert T(\rho _1)-T(\rho _2)\rVert _{\operatorname{tr}} \le (1-\varepsilon )\, \lVert \rho _1-\rho _2\rVert _{\operatorname{tr}}. \label{eq:choi_doeblin_contraction} \end{align}

See [ Wol12 , Chapter 8, Eq. (8.86) ] .

Proof

Since \(\operatorname{tr}[Y]=1\), \(T'_Y\) is trace-preserving: \(\operatorname{tr}[T'_Y(X)]=\operatorname{tr}[X]\operatorname{tr}[Y]=\operatorname{tr}[X]\). Since \(Y\) is Hermitian and \(\operatorname{tr}[X]\) is real for Hermitian \(X\), \(T'_Y\) is Hermiticity-preserving. By Lemma 13.1.31, \(\tau (T'_Y)=(1/D)(Y\otimes \mathbb {1})\). Linearity of the Choi assignment gives \(\tau (T-\varepsilon T'_Y)=\tau (T)-\varepsilon \tau (T'_Y) \ge 0\), so \(T-\varepsilon T'_Y\) is completely positive (the rectangular Choi CP equivalence, Theorem 2.6.6) and hence positive. Theorem 13.1.29 then yields the asserted bound.

13.2 Projective pinching

Theorem 13.2.1 Projective pinching inequality

Let \(\rho \in M_{d}(\mathbb {C})\) be a density matrix and let \(P_1,\ldots ,P_k\in M_{d}(\mathbb {C})\) be orthogonal projections satisfying \(\sum _{i=1}^k P_i=\mathbb {1}\). Define

\begin{align} T(X) & = \sum _{i=1}^k P_iXP_i . \end{align}

Then

\begin{align} \rho & \le kT(\rho ). \label{eq:projective_pinching} \end{align}

This is [ Wol12 , Chapter 8, Eq. (8.56) ] .

Proof

For \(A_{ij}=P_i-P_j\), positivity of \(\rho \) gives \(A_{ij}\rho A_{ij}^{\dagger }\ge 0\). Since the projections are Hermitian and sum to the identity, expansion and collection of the four double sums gives

\begin{align} \sum _{i,j=1}^k A_{ij}\rho A_{ij}^{\dagger } & = 2\bigl(kT(\rho )-\rho \bigr). \end{align}

Thus \(kT(\rho )-\rho \) is positive semidefinite.

13.3 Birkhoff’s theorem

Birkhoff’s theorem characterizes the convex geometry of doubly stochastic matrices.

Theorem 13.3.1 Birkhoff (doubly-stochastic part)

The set of doubly stochastic matrices in \(M_d(\mathbb {R})\) is the convex hull of the \(d\times d\) permutation matrices, and its extreme points are exactly the permutation matrices. See [ Wol12 , Chapter 8, Theorem 8.6 ] . This statement concerns only doubly stochastic matrices; the corresponding assertion of Theorem 8.6 for doubly substochastic matrices is not included here.

Proof

Let \(\mathcal{DS}_d \subset M_d(\mathbb {R})\) be the set of \(d \times d\) doubly stochastic matrices and let \(\mathcal{P}_d\) be the set of \(d \times d\) permutation matrices. The Birkhoff–von Neumann decomposition gives

\begin{align} \mathcal{DS}_d & = \operatorname {conv}(\mathcal{P}_d), \\ \operatorname {ext}(\mathcal{DS}_d) & = \mathcal{P}_d . \end{align}

Indeed, every doubly stochastic matrix is a convex combination of permutation matrices, while convexity of \(\mathcal{DS}_d\) gives the converse inclusion in the first identity. Suppose that a permutation matrix \(P_\sigma \) is a convex combination \(P_\sigma =tA+(1-t)B\), where \(0{\lt}t{\lt}1\) and \(A,B\) are doubly stochastic. Nonnegativity and the row sums give

\begin{align} (P_\sigma )_{ij}=0 & \Longrightarrow A_{ij}=B_{ij}=0, \\ 1=\sum _j A_{ij} & = A_{i,\sigma (i)}, \\ 1=\sum _j B_{ij} & = B_{i,\sigma (i)}. \end{align}

Thus \(A=B=P_\sigma \), so every permutation matrix is extreme. Conversely, a non-permutation matrix has a Birkhoff decomposition involving at least two distinct permutation matrices and is therefore not extreme, which proves the second identity.

13.4 Von Neumann entropy

Definition 13.4.1 Von Neumann entropy
#

For a Hermitian matrix \(\rho \in M_{D}(\mathbb {C})\) with eigenvalues \(\lambda _0,\ldots ,\lambda _{D-1}\), the von Neumann entropy is

\begin{align} S(\rho ) & =-\sum _{i=0}^{D-1}\lambda _i\log \lambda _i, \label{eq:entropy_definition} \end{align}

where \(0\log 0:=0\).

Lemma 13.4.2 Entropy respects equality
#

If two Hermitian matrices are equal, then their von Neumann entropies are equal.

Proof

Substitute the equality of the matrices in the defining eigenvalue sum.

Lemma 13.4.3 Entropy of the zero matrix
#

The zero matrix has zero von Neumann entropy: \(S(0)=0\).

Proof

All eigenvalues of the zero matrix are zero, and \(0\log 0=0\).

Theorem 13.4.4 Entropy is non-negative

For any density matrix \(\rho \), \(S(\rho )\ge 0\).

Proof

Each eigenvalue \(\lambda _i\) of a density matrix satisfies \(0\le \lambda _i\le 1\), and \(-x\log x\ge 0\) on \([0,1]\).

Lemma 13.4.5 Eigenvalue sum
#

The eigenvalues of a density matrix sum to \(1\).

Proof

Follows from \(\operatorname{tr}(\rho )=\sum _i\lambda _i=1\).

Theorem 13.4.6 Eigenvalue bound
#

Each eigenvalue of a density matrix lies in \([0,1]\).

Proof

Non-negativity comes from positive semidefiniteness. The upper bound follows because the eigenvalues are non-negative and sum to \(1\).

Theorem 13.4.7 Entropy upper bound

For a density matrix \(\rho \in M_{D}(\mathbb {C})\) with \(D\ge 1\), one has \(S(\rho )\le \log D\).

Proof

By Jensen’s inequality applied to the concave function \(-x\log x\), the entropy is maximized when all eigenvalues equal \(1/D\).

Theorem 13.4.8 Entropy bounded by log of rank

For a density matrix \(\rho \) of rank \(r\), one has \(S(\rho )\le \log r\). This refines the dimension bound: only the nonzero eigenvalues contribute to the entropy, and there are exactly \(r\) of them.

Proof

The entropy is the \(-x\log x\) sum over all eigenvalues; the \(D-r\) zero eigenvalues contribute nothing. Jensen’s inequality applied to \(-x\log x\) over the \(r\) nonzero eigenvalues, with uniform weights \(1/r\), gives \(S(\rho )\le \log r\), the maximum attained when each nonzero eigenvalue equals \(1/r\). The rank equals the number of nonzero eigenvalues of the Hermitian matrix \(\rho \).

Lemma 13.4.9 Entropy from characteristic polynomial roots

The von Neumann entropy of a Hermitian matrix is the \(-x\log x\) sum over the real parts of the roots of its characteristic polynomial \(\chi _\rho \):

\begin{align} S(\rho ) & =\sum _{\lambda \in \mathrm{roots}(\chi _\rho )} ({-}\operatorname{Re}(\lambda )\log \operatorname{Re}(\lambda )). \label{eq:entropy_charpoly} \end{align}
Proof

The eigenvalues of a Hermitian matrix are exactly the roots of its characteristic polynomial, counted with multiplicity, so the eigenvalue sum defining \(S(\rho )\) equals the displayed sum over roots.

Theorem 13.4.10 Trace-logarithm form of the entropy

Let \(\rho \) be a Hermitian matrix and let \(\log \rho \) be its logarithm defined through the functional calculus. Then

\begin{align} S(\rho ) & =-\operatorname{Re}(\operatorname{tr}(\rho \log \rho )). \label{eq:entropy_trace_log} \end{align}
Proof

A Hermitian matrix and its logarithm are simultaneously diagonalized by a unitary \(U\) with \(\rho =U\operatorname{diag}(\lambda _i)U^*\), so

\begin{align} \operatorname{tr}(\rho \log \rho ) & =\operatorname{tr}\! (U\operatorname{diag}(\lambda _i\log \lambda _i)U^*) =\sum _i\lambda _i\log \lambda _i. \notag \end{align}

Its negative is \(\sum _i({-}\lambda _i\log \lambda _i)=S(\rho )\). Zero eigenvalues contribute nothing under the convention \(0\log 0=0\), so no full-support assumption is required.

Remark 13.4.11
#

The logarithm is the totalized real logarithm, with \(\log x=\log \lvert x\rvert \) and \(\log 0=0\); both sides of the identity use it on every eigenvalue, so the equality holds for an arbitrary Hermitian matrix. It coincides with the physical entropy \(-\operatorname{tr}(\rho \log \rho )\) precisely when \(\rho \) is positive semidefinite. In that case (a density matrix \(\rho \), positive semidefinite with unit trace) the eigenvalues \(\lambda _i\) are non-negative, \(\rho \log \rho \) is Hermitian with real trace, and the real-part extraction is superfluous: the identity reduces to the standard expression \(S(\rho )=-\operatorname{tr}(\rho \log \rho )\).

Definition 13.4.12 Quantum relative entropy
#

For matrices \(\rho ,\sigma \in M_{D}(\mathbb {C})\), define the trace-log expression

\begin{align} D(\rho \Vert \sigma ) & =\operatorname{Re}\operatorname{tr}(\rho (\log \rho -\log \sigma )). \label{eq:entropy_relative_definition} \end{align}

On the physical domain where \(\rho \) is a density matrix and \(\sigma \) is positive definite, this is the Umegaki relative entropy.

Lemma 13.4.13 Trace-log splitting for relative entropy

For matrices \(\rho ,\sigma \in M_{D}(\mathbb {C})\),

\begin{align} D(\rho \Vert \sigma ) & =\operatorname{Re}\operatorname{tr}(\rho \log \rho )-\operatorname{Re}\operatorname{tr}(\rho \log \sigma ). \label{eq:entropy_relative_split} \end{align}
Proof

Expand the matrix product over the difference \(\log \rho -\log \sigma \) and use linearity of the trace and of the real part.

Lemma 13.4.14 Relative entropy with identical arguments
#

For every matrix \(\rho \in M_{D}(\mathbb {C})\), one has \(D(\rho \Vert \rho )=0\).

Proof

The logarithmic difference \(\log \rho -\log \rho \) vanishes.

Lemma 13.4.15 Relative entropy with zero first argument
#

For every matrix \(\sigma \in M_{D}(\mathbb {C})\), one has \(D(0\Vert \sigma )=0\).

Proof

The trace-log expression is multiplied on the left by the zero matrix.

Theorem 13.4.16 Relative entropy in entropy form

If \(\rho \) is Hermitian, then

\begin{align} D(\rho \Vert \sigma ) & =-S(\rho )-\operatorname{Re}\operatorname{tr}(\rho \log \sigma ). \label{eq:entropy_relative_entropy_form} \end{align}
Proof

Apply Lemma 13.4.13 to write

\begin{align} D(\rho \Vert \sigma ) & =\operatorname{Re}\operatorname{tr}(\rho \log \rho )-\operatorname{Re}\operatorname{tr}(\rho \log \sigma ). \notag \end{align}

The trace-logarithm identity \(\operatorname{Re}\operatorname{tr}(\rho \log \rho )=-S(\rho )\) then gives the result.

Theorem 13.4.17 Klein’s inequality

Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices with \(\sigma \) of full rank. Then the relative entropy is non-negative, \(D(\rho \Vert \sigma )\ge 0\).

Proof

Diagonalize \(\rho =\sum _i p_i|e_i\rangle \! \langle e_i|\) and \(\sigma =\sum _j q_j|f_j\rangle \! \langle f_j|\) in their eigenbases, with all \(q_j{\gt}0\) since \(\sigma \) has full rank. The overlap numbers \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\) are non-negative with row sums and column sums equal to \(1\), because the two eigenbases are orthonormal. A trace computation gives

\begin{align} \operatorname{Re}\operatorname{tr}(\rho \log \rho ) & =\sum _i p_i\log p_i, \notag \\ \operatorname{Re}\operatorname{tr}(\rho \log \sigma ) & =\sum _{i,j}p_iP_{ij}\log q_j, \notag \\ D(\rho \Vert \sigma ) & =\sum _i p_i\log p_i-\sum _{i,j}p_iP_{ij}\log q_j. \notag \end{align}

The row and column sums, together with the trace-one normalizations, give \(\sum _{i,j}P_{ij}(p_i-q_j)=\sum _i p_i-\sum _jq_j=0\). Hence

\begin{align} D(\rho \Vert \sigma ) & =\sum _{i,j}P_{ij} \bigl[p_i\log p_i-p_i\log q_j-(p_i-q_j)\bigr]. \notag \end{align}

Each bracket is non-negative. If \(p_i{\gt}0\), put \(x=q_j/p_i{\gt}0\); the bracket is \(p_i(x-1-\log x)\geq 0\) by \(\log x\leq x-1\). If \(p_i=0\), the bracket equals \(q_j{\gt}0\). Therefore \(D(\rho \Vert \sigma )\geq 0\).

Theorem 13.4.18 Klein’s inequality, support form

Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices satisfying the support condition \(\ker \sigma \subseteq \ker \rho \), that is, every vector annihilated by \(\sigma \) is annihilated by \(\rho \). Then the relative entropy is non-negative, \(D(\rho \Vert \sigma )\ge 0\).

Proof

Regularize \(\sigma \) by the trace-one perturbation \(\sigma _\varepsilon '=(1+\varepsilon D)^{-1}(\sigma +\varepsilon \mathbb {1})\), which is positive definite for every \(\varepsilon {\gt}0\), hence of full rank. The full-rank Klein inequality (Theorem 13.4.17) gives \(D(\rho \Vert \sigma _\varepsilon ')\ge 0\). The perturbation shares the eigenbasis \(\{ |f_j\rangle \} \) of \(\sigma \), so the cross term reduces to a scalar sum over the eigenvalues \(q_j\) of \(\sigma \),

\begin{align} \operatorname{Re}\operatorname{tr}(\rho \log \sigma _\varepsilon ’) & =\sum _j\langle f_j|\rho |f_j\rangle \log \! ((1+\varepsilon D)^{-1}(q_j+\varepsilon )). \notag \end{align}

For \(q_j{\gt}0\) the scalar logarithm converges to \(\log q_j\); for \(q_j=0\) the eigenvector \(|f_j\rangle \) lies in \(\ker \sigma \subseteq \ker \rho \), so its diagonal weight \(\langle f_j|\rho |f_j\rangle \) vanishes and the summand is identically zero. No eigenvalue-continuity input is needed, so \(D(\rho \Vert \sigma _\varepsilon ')\to D(\rho \Vert \sigma )\) as \(\varepsilon \to 0^+\) and the inequality passes to the limit.

Theorem 13.4.19 Equality case of Klein’s inequality

Let \(\rho ,\sigma \in M_{D}(\mathbb {C})\) be density matrices with \(\sigma \) of full rank. Then the relative entropy vanishes exactly when the states coincide, \(D(\rho \Vert \sigma )=0\iff \rho =\sigma \). Together with nonnegativity, this is the order property that makes \(D\) a divergence.

Proof

If \(\rho =\sigma \) then \(\log \rho -\log \sigma =0\) and \(D(\rho \Vert \sigma )=0\). For the converse, diagonalize \(\rho =\sum _i p_i|e_i\rangle \! \langle e_i|\) and \(\sigma =\sum _j q_j|f_j\rangle \! \langle f_j|\), with all \(q_j{\gt}0\) by full rank, and set \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\). As in Theorem 13.4.17,

\begin{align} D(\rho \Vert \sigma ) & =\sum _{i,j}P_{ij} [p_i\log p_i-p_i\log q_j-(p_i-q_j)], \notag \end{align}

a sum of non-negative terms: each is bounded below by the tangent inequality \(\log x\le x-1\) at \(x=q_j/p_i\), while the linear remainder telescopes through the doubly stochastic row and column sums, \(\sum _{i,j}P_{ij}(p_i-q_j)=\sum _ip_i-\sum _jq_j=0\). If the total vanishes then every term vanishes. On a row with \(p_i=0\) the term reads \(P_{ij}q_j\), so \(P_{ij}=0\); on a row with \(p_i{\gt}0\) the term forces the tangent inequality to be tight, \(\log (q_j/p_i)=q_j/p_i-1\), and strict concavity of the logarithm (\(\log x{\lt}x-1\) for \(x\neq 1\)) gives \(q_j=p_i\). Hence \(q_j=p_i\) whenever \(\langle e_i | f_j \rangle \neq 0\). Writing \(W=U_\rho ^\dagger U_\sigma \) for the overlap of the eigenvector unitaries, this matching says \(W\operatorname{diag}(q)=\operatorname{diag}(p)W\). Thus, conjugating \(\sigma \) into the eigenbasis of \(\rho \) and using the unitarity \(WW^\dagger =1\),

\begin{align} U_\rho ^\dagger \sigma U_\rho & =W\operatorname{diag}(q)W^\dagger =\operatorname{diag}(p)WW^\dagger =\operatorname{diag}(p), \notag \end{align}

the spectral diagonal of \(\rho \). Therefore \(\sigma =U_\rho \operatorname{diag}(p)U_\rho ^\dagger =\rho \).

Theorem 13.4.20 Joint convexity of the relative entropy

On pairs of positive definite matrices in \(M_{D}(\mathbb {C})\), the map \((\rho ,\sigma )\mapsto D(\rho \Vert \sigma )\) is jointly convex.

Proof

For \(s\in [0,1)\) consider the approximant

\begin{align} g_s(\rho ,\sigma ) & =\frac{\operatorname{tr}\rho -\operatorname{Re}\operatorname{tr}(\rho ^s\sigma ^{1-s})}{1-s}. \label{eq:entropy_convex_approximant} \end{align}

The trace \(\operatorname{tr}\rho \) is real-affine in the pair, hence convex, while \((\rho ,\sigma )\mapsto \operatorname{Re}\operatorname{tr}(\rho ^s\sigma ^{1-s})\) is jointly concave by the \(K=\mathbb {1}\) case of the Lieb concavity theorem (Corollary 7.7.14). Their difference, scaled by the non-negative factor \((1-s)^{-1}\), is therefore jointly convex, so each \(g_s\) is convex.

Writing \(\rho \) and \(\sigma \) in their eigenbases with eigenvalues \(p_i,q_j{\gt}0\) and overlap weights \(P_{ij}=\lvert \langle e_i | f_j \rangle \rvert ^2\), the approximant becomes the double sum \(g_s(\rho ,\sigma )=\sum _{i,j}(1-s)^{-1}(p_i-p_i^sq_j^{1-s})P_{ij}\). As \(s\to 1^-\) each per-pair term converges, via the scalar limit \((c^u-1)/u\to \log c\), to \(p_i(\log p_i-\log q_j)P_{ij}\), whose sum equals \(D(\rho \Vert \sigma )\). Thus \(g_s\to D\) pointwise on positive definite pairs, and since the pointwise limit of convex functions is convex, \(D\) is jointly convex.

Theorem 13.4.21 Joint convexity on the support domain

On pairs of density matrices \((\rho ,\sigma )\) in \(M_{D}(\mathbb {C})\) satisfying the support condition \(\ker \sigma \subseteq \ker \rho \), the map \((\rho ,\sigma )\mapsto D(\rho \Vert \sigma )\) is jointly convex.

Proof

The domain is convex: for a strict convex combination of two such pairs the kernel of \(a\sigma _1+b\sigma _2\) is \(\ker \sigma _1\cap \ker \sigma _2\), since the quadratic forms of the positive semidefinite summands are non-negative and their positively weighted sum vanishes only when each does, and a positive semidefinite matrix annihilates exactly the vectors of zero quadratic form; the two pointwise support inclusions then give \(\ker \sigma _1\cap \ker \sigma _2\subseteq \ker \rho _1\cap \ker \rho _2\).

Regularize both arguments through the affine trace-one perturbation \(M_\varepsilon =(1+\varepsilon N)^{-1}(M+\varepsilon \mathbb {1})\), where \(N\) is the matrix dimension, which is positive definite for every \(\varepsilon {\gt}0\). Because the perturbation is affine, it commutes with the convex combination, so the four regularized endpoints and the regularized mixture all lie in the positive definite domain and the positive definite joint convexity (Theorem 13.4.20) gives the two-point inequality for the regularized pairs. As \(\varepsilon \to 0^+\), the regularized relative entropy converges on each pair of the support domain: \(D(\rho _\varepsilon \Vert \sigma _\varepsilon ) \to D(\rho \Vert \sigma )\). Indeed the perturbation shares the eigenbasis of its argument, so each trace-logarithm term is the diagonal sum \(\sum _jw_j(\varepsilon ) \log ((1+\varepsilon N)^{-1}(q_j+\varepsilon ))\), where \(q_j\) runs over the eigenvalues of \(\sigma \) and \(w_j(\varepsilon )\) is the corresponding diagonal weight. For \(q_j{\gt}0\), the scalar factor converges to \(\log q_j\), while at a zero eigenvalue the support condition makes the weight vanish, with \(w_j(\varepsilon )=(1+\varepsilon N)^{-1}\varepsilon \), so the summand is

\begin{align} & (1+\varepsilon N)^{-1}\varepsilon \log ((1+\varepsilon N)^{-1}\varepsilon )\longrightarrow 0 \notag \end{align}

by \(x\log x\to 0\) as \(x\to 0^+\). The two-point inequality therefore passes to the limit, giving joint convexity on the support domain.

Lemma 13.4.22 Functional calculus under unitary conjugation
#

For a Hermitian matrix \(A\), a real function \(f\), and a unitary \(U\), the continuous functional calculus satisfies

\begin{align} f(UAU^\dagger ) & =Uf(A)U^\dagger . \label{eq:entropy_cfc_conjugation} \end{align}
Proof

The conjugation \(\varphi :x\mapsto UxU^\dagger \) is a star-algebra automorphism of the matrix algebra, and the continuous functional calculus commutes with such automorphisms, \(\varphi (f(A))=f(\varphi (A))\); since \(\varphi (A)=UAU^\dagger \), this is the claim.

Lemma 13.4.23 Matrix logarithm under unitary conjugation
#

For a Hermitian matrix \(A\) and a unitary \(U\),

\begin{align} \log (UAU^\dagger ) & =U(\log A)U^\dagger . \label{eq:entropy_log_conjugation} \end{align}
Proof

This is the \(f=\log \) case of Lemma 13.4.22: \(f(UAU^\dagger )=Uf(A)U^\dagger \) specialized to the real logarithm gives \(\log (UAU^\dagger )=U(\log A)U^\dagger \).

Theorem 13.4.24 Unitary invariance of the quantum relative entropy

For Hermitian matrices \(\rho ,\sigma \) and a unitary \(U\),

\begin{align} D(U\rho U^\dagger \Vert U\sigma U^\dagger ) & =D(\rho \Vert \sigma ). \label{eq:entropy_unitary_invariance} \end{align}
Proof

Carry the two logarithms through the conjugation by Lemma 13.4.23, so that

\begin{align} D(U\rho U^\dagger \Vert U\sigma U^\dagger ) & =\operatorname{Re}\operatorname{tr}\! (U\rho (\log \rho -\log \sigma )U^\dagger ) =\operatorname{Re}\operatorname{tr}\! (\rho (\log \rho -\log \sigma )) =D(\rho \Vert \sigma ), \notag \end{align}

where the middle equality is trace cyclicity together with \(U^\dagger U=1\).

Lemma 13.4.25 Logarithm of a tensor product of positive definite matrices
#

For positive definite matrices \(\rho \) and \(\tau \),

\begin{align} \log (\rho \otimes \tau ) & =\log \rho \otimes \mathbb {1}+\mathbb {1}\otimes \log \tau . \label{eq:entropy_log_tensor} \end{align}
Proof

The tensor product factors as \(\rho \otimes \tau =(\rho \otimes \mathbb {1})(\mathbb {1}\otimes \tau )\) into a pair of commuting positive definite matrices, so the logarithm of the product is the sum of the logarithms of the factors. Each unital embedding \(A\mapsto A\otimes \mathbb {1}\) and \(B\mapsto \mathbb {1}\otimes B\) is a continuous star-algebra homomorphism, hence commutes with the functional calculus, which carries each logarithm onto its factor: \(\log (\rho \otimes \mathbb {1})=\log \rho \otimes \mathbb {1}\) and \(\log (\mathbb {1}\otimes \tau )=\mathbb {1}\otimes \log \tau \).

Theorem 13.4.26 Ancilla additivity of the quantum relative entropy

For positive definite matrices \(\rho ,\sigma \) and a positive definite matrix \(\tau \) of unit trace,

\begin{align} D(\rho \otimes \tau \Vert \sigma \otimes \tau ) & =D(\rho \Vert \sigma ). \label{eq:entropy_ancilla_additivity} \end{align}
Proof

Splitting each tensor logarithm by Lemma 13.4.25, the common \(\mathbb {1}\otimes \log \tau \) terms cancel in the difference, leaving \(\log (\rho \otimes \tau )-\log (\sigma \otimes \tau ) =(\log \rho -\log \sigma )\otimes \mathbb {1}\). Hence

\begin{align} D(\rho \otimes \tau \Vert \sigma \otimes \tau ) & =\operatorname{Re}\operatorname{tr}\! ((\rho (\log \rho -\log \sigma ))\otimes \tau ) =\operatorname{Re}(\operatorname{tr}(\rho (\log \rho -\log \sigma ))\cdot \operatorname{tr}\tau ), \notag \end{align}

where the trace of a tensor product factors as a product of traces. As \(\operatorname{tr}\tau =1\), the right-hand side is \(D(\rho \Vert \sigma )\).

Lemma 13.4.27 Root-of-unity character-sum orthogonality
#

Let \(\zeta \) be a primitive \(d\)-th root of unity and let \(i,j\) range over the residues modulo \(d\). Then

\begin{align} \sum _{b=0}^{d-1}\zeta ^{bi}\overline{\zeta ^{bj}} & = \begin{cases} d, & i=j, \\ 0, & i\ne j. \end{cases} \label{eq:entropy_root_orthogonality} \end{align}
Proof

Each summand equals \(\xi ^b\) with \(\xi =\zeta ^i\overline{\zeta }^{\, j}\), and \(\xi ^d=1\). When \(i=j\) the base \(\xi \) equals \(1\) and the sum is \(d\). When \(i\ne j\) the base is a root of unity different from \(1\), so the geometric sum \((\xi -1)\sum _b\xi ^b=\xi ^d-1=0\) forces the sum to vanish; the equivalence \(\xi =1\iff i=j\) uses that \(\overline{\zeta }=\zeta ^{-1}\) and the injectivity of \(b\mapsto \zeta ^b\) on residues.

Theorem 13.4.28 Unitarity of the Weyl operators

Fix a dimension \(d\ge 1\) and a primitive \(d\)-th root of unity \(\zeta \). The cyclic shift \(X\), the clock operator \(Z\), and every Weyl operator \(W(a,b)=X^aZ^b\) are unitary.

Proof

The cyclic shift is the permutation matrix of a cyclic permutation. The clock operator is diagonal, and each diagonal entry is a power of \(\zeta \), hence has modulus one. Thus \(X\) and \(Z\) are unitary, and so is every product \(X^aZ^b\).

Theorem 13.4.29 Weyl-operator unitary 1-design twirl

Fix a dimension \(d\ge 1\) and a primitive \(d\)-th root of unity \(\zeta \), and let \(X\) be the cyclic shift \(|i\rangle \mapsto |i+1\rangle \) and \(Z=\operatorname{diag}(\zeta ^0,\ldots ,\zeta ^{d-1})\) the clock operator. For every matrix \(M\) on \(\mathbb {C}^d\), the uniform average of the conjugations by the \(d^2\) Weyl operators \(W(a,b)=X^aZ^b\) is the completely depolarizing channel:

\begin{align} \frac{1}{d^2}\sum _{a,b=0}^{d-1} W(a,b)MW(a,b)^\dagger & =\frac{\operatorname{tr}M}{d}\mathbb {1}. \label{eq:entropy_weyl_twirl} \end{align}
Proof

The double average factors into a clock average followed by a shift average. The clock average \(\sum _bZ^bM(Z^b)^\dagger \) multiplies the entry \(M_{ij}\) by \(\sum _b\zeta ^{bi}\overline{\zeta ^{bj}}\), which by Lemma 13.4.27 is \(d\) when \(i=j\) and \(0\) otherwise; the result is \(d\) times the diagonal part of \(M\). The shift average \(\sum _aX^a(\operatorname{diag}v)(X^a)^\dagger \) cyclically permutes the diagonal entries, so each diagonal position receives the full sum \(\sum _kv_k=\operatorname{tr}M\), giving \((\operatorname{tr}M)\mathbb {1}\). Combining the two factors of \(d\) with the prefactor \(d^{-2}\) leaves \((\operatorname{tr}M/d)\mathbb {1}\).

Lemma 13.4.30 Twirl as partial trace tensored with the maximally mixed ancilla

Fix a dimension \(d_C\ge 1\) and a primitive \(d_C\)-th root of unity \(\zeta \). For every matrix \(M\) on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\), the uniform average of the conjugations by the \(d_C^2\) unitaries \(\mathbb {1}_S\otimes W(a,b)\) on the second factor is the partial trace over that factor tensored with the maximally mixed state \(\mathbb {1}_C/d_C\):

\begin{align} \frac{1}{d_C^2}\sum _{a,b} (\mathbb {1}_S\otimes W(a,b))M(\mathbb {1}_S\otimes W(a,b))^\dagger & =(\operatorname{tr}_C M)\otimes (\mathbb {1}_C/d_C). \label{eq:entropy_partial_trace_twirl} \end{align}
Proof

On each pair of blocks indexed by the first factor, the conjugation by \(\mathbb {1}_S\otimes W(a,b)\) acts as the Weyl conjugation \(W(a,b)(\cdot )W(a,b)^\dagger \) of the corresponding block of \(M\). Averaging over the \(d_C^2\) Weyl operators sends each block to the depolarizing channel by Theorem 13.4.29, replacing it by its trace times \(\mathbb {1}_C/d_C\). The block trace is exactly the corresponding entry of the partial trace \(\operatorname{tr}_C M\), so the average is \((\operatorname{tr}_C M)\otimes (\mathbb {1}_C/d_C)\).

Theorem 13.4.31 The maximally mixed ancilla

For \(d_C\ge 1\), the matrix \(\tau _C=d_C^{-1}\mathbb {1}_C\) is positive definite and has trace one.

Proof

The identity is positive definite and \(d_C^{-1}{\gt}0\), so \(\tau _C\) is positive definite. Moreover, \(\operatorname{tr}\tau _C=d_C^{-1}\operatorname{tr}\mathbb {1}_C=1\).

Theorem 13.4.32 Covariance of the matrix logarithm under reindexing
#

Let \(A\) be a Hermitian matrix on a finite index set, and let \(e\) be a bijection onto another finite index set. Then

\begin{align} \log (A_{e^{-1},e^{-1}}) & =(\log A)_{e^{-1},e^{-1}}. \notag \end{align}
Proof

Reindexing is a star-algebra isomorphism and therefore commutes with the continuous functional calculus for the real logarithm.

Lemma 13.4.33 Reindexing invariance of the quantum relative entropy

For Hermitian matrices \(\rho ,\sigma \) on a finite index set and any bijection \(e\) from that set onto another finite set,

\begin{align} D(\rho _{e^{-1},e^{-1}}\Vert \sigma _{e^{-1},e^{-1}}) & =D(\rho \Vert \sigma ). \label{eq:entropy_reindexing} \end{align}
Proof

Write \(D(\rho \Vert \sigma ) =\operatorname{Re}\operatorname{tr}(\rho (\log \rho -\log \sigma ))\). The matrix logarithm is covariant under reindexing, \(\log (\rho _{e^{-1},e^{-1}}) =(\log \rho )_{e^{-1},e^{-1}}\), because reindexing is a star-algebra isomorphism and so commutes with the continuous functional calculus. The same isomorphism preserves products and the trace,

\begin{align} (\rho M)_{e^{-1},e^{-1}} & =\rho _{e^{-1},e^{-1}}M_{e^{-1},e^{-1}}, \notag \\ \operatorname{tr}(M_{e^{-1},e^{-1}}) & =\operatorname{tr}(M). \notag \end{align}

Taking \(M=\log \rho -\log \sigma \) and applying these three identities gives \(D(\rho _{e^{-1},e^{-1}}\Vert \sigma _{e^{-1},e^{-1}}) =D(\rho \Vert \sigma )\).

For positive definite matrices \(\rho ,\sigma \) on a tensor product of a system factor and an ancilla factor,

\begin{align} D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma ) & \le D(\rho \Vert \sigma ), \label{eq:entropy_data_processing} \end{align}

where \(\operatorname{tr}_C\) is the partial trace over the ancilla factor. This is the positive-definite base case; the source inequality on the support domain \(\ker \sigma \subseteq \ker \rho \) is Theorem 13.4.35.

Proof

By ancilla additivity (Theorem 13.4.26) the reduced-state relative entropy equals \(D\bigl((\operatorname{tr}_C\rho )\otimes (\mathbb {1}_C/d_C)\Vert (\operatorname{tr}_C\sigma )\otimes (\mathbb {1}_C/d_C)\bigr)\). By Lemma 13.4.30 each tensored reduced state is the uniform average of the conjugations by the \(d_C^2\) unitaries \(U_{ab}=\mathbb {1}_S\otimes W(a,b)\), so this is the relative entropy of a convex combination of the conjugated pairs \((U_{ab}\rho U_{ab}^\dagger ,U_{ab}\sigma U_{ab}^\dagger )\) with equal weights \(d_C^{-2}\). Joint convexity bounds it above by the same convex combination of the per-term relative entropies, each of which equals \(D(\rho \Vert \sigma )\) by unitary invariance (Theorem 13.4.24). As the weights sum to one, the bound is \(D(\rho \Vert \sigma )\).

Theorem 13.4.35 Data-processing inequality on the support domain

For positive semidefinite \(\rho ,\sigma \) on a tensor product of a system factor and an ancilla factor of dimension \(d_C\), with the support condition \(\ker \sigma \subseteq \ker \rho \),

\begin{align} D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma ) & \le D(\rho \Vert \sigma ), \label{eq:entropy_data_processing_support} \end{align}

where \(\operatorname{tr}_C\) is the partial trace over the ancilla factor.

Proof

Let \(N\) be the dimension of the joint space. Regularize both arguments through the affine trace-shrinking perturbation \(M_\varepsilon =(1+N\varepsilon )^{-1}(M+\varepsilon \mathbb {1})\), which is positive definite for every \(\varepsilon {\gt}0\), so the positive-definite data-processing inequality (Theorem 13.4.34) gives

\begin{align} D(\operatorname{tr}_C\rho _\varepsilon \Vert \operatorname{tr}_C\sigma _\varepsilon ) & \le D(\rho _\varepsilon \Vert \sigma _\varepsilon ). \notag \end{align}

The right-hand side converges to \(D(\rho \Vert \sigma )\) as \(\varepsilon \to 0^+\), by the same shared-eigenbasis scalar-limit argument as the joint convexity on the support domain (Theorem 13.4.21).

For the left-hand side, the partial trace of the regularization is a differently scaled regularization of the marginal: because the partial trace is linear and \(\operatorname{tr}_C\mathbb {1}=d_C\mathbb {1}\),

\begin{align} \operatorname{tr}_C((1+N\varepsilon )^{-1}(M+\varepsilon \mathbb {1})) & =(1+N\varepsilon )^{-1} (\operatorname{tr}_C M+d_C\varepsilon \mathbb {1}). \notag \end{align}

This is the affine regularization of \(\operatorname{tr}_C M\) with the same scaling rate \(N\) but shift rate \(d_C\) and the smaller identity on the system factor. The support condition transfers to the marginals, \(\ker (\operatorname{tr}_C\sigma )\subseteq \ker (\operatorname{tr}_C\rho )\): a vector annihilated by \(\operatorname{tr}_C\sigma \) has vanishing marginal quadratic form, which splits into the non-negative joint quadratic forms of its single-ancilla lifts, so each lift lies in \(\ker \sigma \), hence in \(\ker \rho \), and reassembling the ancilla sum shows the vector lies in \(\ker (\operatorname{tr}_C\rho )\). With this support condition the arbitrary-rate affine regularization has the same both-arguments limit:

\begin{align} D(\operatorname{tr}_C\rho _\varepsilon \Vert \operatorname{tr}_C\sigma _\varepsilon ) & \longrightarrow D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma ) \quad \text{as }\varepsilon \to 0^+. \notag \end{align}

Passing the inequality through the two limits gives the support-domain bound.

Definition 13.4.36 Inverse square root on the support
#

Let \(\tau =\sum _i\lambda _i|i\rangle \! \langle i|\) be positive semidefinite. Its inverse square root on the support is

\begin{align} \tau ^{-1/2}_{\mathrm{supp}} & =\sum _{\lambda _i{\gt}0}\lambda _i^{-1/2}|i\rangle \! \langle i|. \label{eq:entropy_support_inv_sqrt} \end{align}
Lemma 13.4.37 Hermiticity of the support inverse square root

The inverse square root of a positive semidefinite matrix on its support is Hermitian.

Lemma 13.4.38 Positivity of the support inverse square root

The inverse square root of a positive semidefinite matrix on its support is positive semidefinite.

Proof

On the non-negative spectrum of \(\tau \), the defining function satisfies

\begin{align} f(x) & = \begin{cases} x^{-1/2}, & x{\gt}0,\\ 0, & x=0 \end{cases} \geq 0 \quad (x\geq 0). \notag \end{align}

The spectral functional calculus therefore gives \(f(\tau )\geq 0\).

Lemma 13.4.39 Support inverse square root of a positive-definite matrix

If \(\tau \) is positive definite, then its inverse square root on the support is its ordinary inverse square root: \(\tau ^{-1/2}_{\mathrm{supp}}=(\sqrt\tau )^{-1}\).

Proof

Every eigenvalue of \(\tau \) is strictly positive. Hence the function defining the support inverse square root agrees on the spectrum with the reciprocal of the positive square-root function.

Lemma 13.4.40 Support inverse-square-root identity

If \(P_\tau \) is the orthogonal projector onto the support of a positive semidefinite matrix \(\tau \), then

\begin{align} \tau ^{-1/2}_{\mathrm{supp}} \tau \tau ^{-1/2}_{\mathrm{supp}} & =P_\tau . \label{eq:entropy_support_sandwich} \end{align}
Proof

In an eigenbasis of \(\tau \), the left-hand side has eigenvalue zero when \(\lambda _i=0\) and eigenvalue \(\lambda _i^{-1/2}\lambda _i\lambda _i^{-1/2}=1\) otherwise.

Lemma 13.4.41 Support generalized-inverse identity

If \(P_\tau \) is the orthogonal projector onto the support of a positive semidefinite matrix \(\tau \), then

\begin{align} \left(\tau ^{-1/2}_{\mathrm{supp}}\right)^2\tau =\tau \left(\tau ^{-1/2}_{\mathrm{supp}}\right)^2 & =P_\tau . \label{eq:entropy_support_generalized_inv} \end{align}
Proof

The support inverse square root commutes with \(\tau \) by spectral functional calculus, so

\begin{align} \left(\tau ^{-1/2}_{\mathrm{supp}}\right)^2\tau & =\tau ^{-1/2}_{\mathrm{supp}} \tau \tau ^{-1/2}_{\mathrm{supp}} =P_\tau , \notag \end{align}

where the last step is Lemma 13.4.40. Taking adjoints gives the identity with \(\tau \) on the left.

Theorem 13.4.42 Multiplicative functional calculus on positive tensor products
#

Let \(A\) and \(B\) be positive semidefinite, and let \(f\colon \mathbb {R}\to \mathbb {R}\) be multiplicative on the non-negative reals. Then

\begin{align} f(A\otimes B)& =f(A)\otimes f(B). \label{eq:entropy_cfc_tensor_product} \end{align}
Proof

Diagonalize \(A\) and \(B\). In the resulting product eigenbasis, the eigenvalues of \(A\otimes B\) are \(a_i b_j\) with \(a_i,b_j\geq 0\), and the claim follows from \(f(a_i b_j)=f(a_i)f(b_j)\).

Let \(A\) and \(B\) be positive semidefinite. Their positive square roots, support inverse square roots, and support projections satisfy

\begin{align} \sqrt{A\otimes B} & =\sqrt A\otimes \sqrt B, \notag \\ (A\otimes B)^{-1/2}_{\mathrm{supp}} & =A^{-1/2}_{\mathrm{supp}}\otimes B^{-1/2}_{\mathrm{supp}}, \notag \\ P_{A\otimes B} & =P_A\otimes P_B. \label{eq:entropy_support_tensor_product} \end{align}

Moreover,

\begin{align} \sqrt A\, A^{-1/2}_{\mathrm{supp}} =A^{-1/2}_{\mathrm{supp}}\sqrt A & =P_A. \label{eq:entropy_support_sqrt_cancellation} \end{align}

The same support projection absorbs the square root:

\begin{align} \sqrt A\, P_A=P_A\sqrt A& =\sqrt A. \notag \end{align}

If \(A\) is positive definite, then \(P_A=\mathbf1\).

Proof

Diagonalize both factors. The eigenvalues of \(A\otimes B\) are the products \(a_i b_j\). Both the square-root function and the function that equals \(x^{-1/2}\) for \(x{\gt}0\) and zero at \(x=0\) are multiplicative on the non-negative reals. For the support projectors, apply the sandwich identity to \(A\otimes B\), factor the support inverse square root and matrix products, and apply the sandwich identity to each factor. The cancellation identities follow entrywise. If \(A\) is positive definite, then \(A^{-1/2}_{\mathrm{supp}}=(\sqrt A)^{-1}\) and \(\sqrt A\) is invertible. Hence the cancellation identity gives \(P_A=\sqrt A(\sqrt A)^{-1}=\mathbf1\).

Let \(\tau \) be positive semidefinite and let \(c{\gt}0\). Then

\begin{align} (c\tau )^{-1/2}_{\mathrm{supp}} & =c^{-1/2}\tau ^{-1/2}_{\mathrm{supp}}. \label{eq:entropy_support_scaling} \end{align}

For every finite-dimensional auxiliary space \(\mathcal H_R\),

\begin{align} (\tau \otimes \mathbf1_R)^{-1/2}_{\mathrm{supp}} & =\tau ^{-1/2}_{\mathrm{supp}}\otimes \mathbf1_R. \label{eq:entropy_support_tensor} \end{align}

The corresponding identity for an auxiliary left factor is

\begin{align} (\mathbf1_L\otimes \tau )^{-1/2}_{\mathrm{supp}} & =\mathbf1_L\otimes \tau ^{-1/2}_{\mathrm{supp}}. \label{eq:entropy_support_tensor_left} \end{align}

Consequently, for \(d_R{\gt}0\),

\begin{align} (\tau \otimes d_R^{-1}\mathbf1_R)^{-1/2}_{\mathrm{supp}} & =\sqrt{d_R} (\tau ^{-1/2}_{\mathrm{supp}}\otimes \mathbf1_R). \label{eq:entropy_support_mixed} \end{align}
Proof

The scalar identity follows from the functional calculus and \(\sqrt{cx}=\sqrt c\sqrt x\) for \(c{\gt}0\). For the support-inverse function \(f(x)=x^{-1/2}\) when \(x{\gt}0\) and \(f(0)=0\), the functional calculus gives

\begin{align} f(\tau \otimes \mathbf1_R) & =f(\tau )\otimes \mathbf1_R. \notag \end{align}

This proves (61); the same argument with the identity as the left factor proves (62). Applying (60) with \(c=d_R^{-1}\) to \(\tau \otimes \mathbf1_R\) then gives (63).

Definition 13.4.45 Support Petz transpose map for a partial trace
#

Let \(\sigma \) be positive semidefinite on \(H_L\otimes H_R\), and set \(\tau =\operatorname{tr}_R\sigma \). The Petz transpose formula on the support of \(\tau \) is

\begin{align} \mathcal R_\sigma (X) & =\sqrt\sigma \bigl(\tau ^{-1/2}_{\mathrm{supp}}X \tau ^{-1/2}_{\mathrm{supp}}\otimes \mathbf1_R\bigr)\sqrt\sigma . \label{eq:entropy_support_petz} \end{align}

This is the support formula of [ HJPW04 , Theorem 3, equation (8) ] . It is not asserted to be trace preserving on operators outside the support of \(\tau \).

Lemma 13.4.46 Support Petz formula

For every matrix \(X\), the support Petz map is given by (64).

Definition 13.4.47 Product reference for a partial trace

Let \(\rho _A\) and \(\rho _{BC}\) be positive semidefinite, and define

\begin{align} \sigma _{ABC} & =\rho _A\otimes \rho _{BC}. \label{eq:entropy_general_product_reference} \end{align}

We use the canonical reassociation from \(A\times (B\times C)\) to \((A\times B)\times C\), so that the right partial trace removes \(C\).

Set \(\rho _B=\operatorname{tr}_C\rho _{BC}\), and let \(P_A\) and \(P_B\) be the support projections of \(\rho _A\) and \(\rho _B\). Then

\begin{align} \operatorname{tr}_C\sigma _{ABC} & =\rho _A\otimes \rho _B, \label{eq:entropy_product_reference_marginal}\\ \sqrt{\sigma _{ABC}} & =\sqrt{\rho _A}\otimes \sqrt{\rho _{BC}}, \notag \\ (\operatorname{tr}_C\sigma _{ABC})^{-1/2}_{\mathrm{supp}} & =\rho _{A,\mathrm{supp}}^{-1/2}\otimes \rho _{B,\mathrm{supp}}^{-1/2}, \notag \\ P_{\operatorname{tr}_C\sigma _{ABC}} & =P_A\otimes P_B. \label{eq:entropy_product_reference_support} \end{align}

Each identity is understood after the same canonical reassociation of the three tensor factors.

Proof

Expand the marginal in a product basis. The remaining identities follow from functional calculus for positive semidefinite tensor products. The support projection is obtained by multiplying the factorized support inverse square root on both sides of the marginal.

Definition 13.4.49 Compression to the first-factor support

If \(P_A\) is the support projection of \(\rho _A\), define

\begin{align} \mathcal S_{P_A}(X) & =P_AXP_A. \label{eq:entropy_support_compression} \end{align}

The operator-Schmidt rank below is a generic bipartite-matrix fact, with no tensor-network content; it is relocated here from the MPDO area-law chapter, whose diagonal-cut and pure-state support-compression bounds cite it across the chapter boundary.

For a bipartite density operator \(\rho \in M_{d_A}(\mathbb {C})\otimes M_{d_B}(\mathbb {C})\), its operator-Schmidt rank is the least integer \(r\) for which there are matrices \(A_t\in M_{d_A}(\mathbb {C})\) and \(B_t\in M_{d_B}(\mathbb {C})\) satisfying

\begin{align} \rho & =\sum _{t=1}^{r}A_t\otimes B_t. \notag \end{align}

No Hermiticity or positivity condition is imposed on the factors. This is the definition in [ DlCDN19 , Equation (1) ] ; the same formula defines the rank of an arbitrary complex bipartite matrix.

For the reference in (65), the raw Petz map for \(\operatorname{tr}_C\) satisfies

\begin{align} \mathcal R_{\sigma _{ABC}} & =\mathcal S_{P_A}\otimes \mathcal R_{\rho _{BC}}. \label{eq:entropy_hjpw_product_support_factorization} \end{align}

Equivalently, for every product operator \(A_0\otimes X_B\),

\begin{align} \mathcal R_{\sigma _{ABC}}(A_0\otimes X_B) & =(P_AA_0P_A)\otimes \mathcal R_{\rho _{BC}}(X_B). \notag \end{align}

The formulas use the canonical identification \((H_A\otimes H_B)\otimes H_C\cong H_A\otimes (H_B\otimes H_C)\).

This is the globally valid ambient-space form of [ HJPW04 , equation (10) ] . The literal identity-tensored formula in that equation is obtained on the support of \(\rho _A\), or after choosing an extension away from that support. For singular \(\rho _A\), the raw map on the full matrix algebra contains the compression \(X\mapsto P_AXP_A\).

Proof

Substitute the tensor factorizations of the square root and marginal support inverse into the Petz sandwich. The first-factor terms reduce to \(P_AA_0P_A\). This proves the formula for product operators. A finite product-operator decomposition proves the linear-map identity.

Let \(X\) be an operator on \(H_A\otimes H_B\). If

\begin{align} (P_A\otimes \mathbf1_B)X(P_A\otimes \mathbf1_B) & =X, \notag \end{align}

then

\begin{align} \mathcal R_{\sigma _{ABC}}(X) & =(\operatorname {id}_A\otimes \mathcal R_{\rho _{BC}})(X). \label{eq:entropy_hjpw_product_supported} \end{align}

If \(\rho _A\) is positive definite, then \(P_A=\mathbf1_A\), and this identity holds for every \(X\).

When \(\rho _A\) is singular, no global identity-tensored formula is asserted for the raw map outside the displayed support. Nor is the generic trace-preserving completion in Definition 13.4.63 asserted to factor: its complementary projection is \(\mathbf1_{AB}-P_A\otimes P_B\), which need not be the identity on \(A\) tensored with a projection on \(B\).

Proof

The support condition makes \(\mathcal S_{P_A}\) act as the identity on \(X\). If \(\rho _A\) is positive definite, its support projection is the identity, so the condition holds on the full matrix algebra.

Definition 13.4.53 Maximally mixed tensor reference
#

Let \(d_A{\gt}0\) and let \(\rho _{BC}\) be positive semidefinite. Define

\begin{align} \sigma _{ABC} & =d_A^{-1}\mathbf1_A\otimes \rho _{BC}. \label{eq:entropy_hjpw_product_reference} \end{align}

Under the canonical reassociation from \(A\times (B\times C)\) to \((A\times B)\times C\), the right partial trace removes \(C\).

For the reference in (71),

\begin{align} \operatorname{tr}_C\sigma _{ABC} & =d_A^{-1}\mathbf1_A\otimes \rho _B, \notag \\ \operatorname{tr}_{AB}\sigma _{ABC} & =\rho _C, \notag \\ \operatorname{tr}(\sigma _{ABC}) & =\operatorname{tr}(\rho _{BC}). \label{eq:entropy_mixed_reference_trace} \end{align}
Proof

Expand the two partial traces in a product basis. The sum over the \(d_A\) diagonal entries cancels the factor \(d_A^{-1}\). Trace invariance under a partial trace then gives the last identity.

Theorem 13.4.55 Support inverse square root of the maximally mixed marginal

If \(\rho _{BC}\) is positive semidefinite and \(\rho _B=\operatorname{tr}_C\rho _{BC}\), then

\begin{align} \sigma _{AB,\mathrm{supp}}^{-1/2} & =\sqrt{d_A}\, \mathbf1_A\otimes \rho _{B,\mathrm{supp}}^{-1/2}. \label{eq:entropy_mixed_reference_support_inverse} \end{align}
Proof

The marginal identity gives \(\sigma _{AB}=d_A^{-1}\mathbf1_A\otimes \rho _B\). Apply the support-inverse scaling identity and the tensor identity for an auxiliary left factor. Since \(d_A{\gt}0\), \((\sqrt{d_A^{-1}})^{-1}=\sqrt{d_A}\).

For the reference in (71), the support Petz map for \(\operatorname{tr}_C\) factors as

\begin{align} \mathcal R_{\sigma _{ABC}} & =\operatorname {id}_A\otimes \mathcal R_{\rho _{BC}}. \label{eq:entropy_hjpw_petz_factorization} \end{align}

after the canonical identification \((H_A\otimes H_B)\otimes H_C\cong H_A\otimes (H_B\otimes H_C)\). This is the maximally mixed specialization of [ HJPW04 , equation (10) ] . It is the tensor-product identity for the raw Petz support formula, not the Hayashi–Koashi–Imoto block decomposition.

Proof

Functional calculus gives \(\sqrt{\sigma _{ABC}} =d_A^{-1/2}\mathbf1_A\otimes \sqrt{\rho _{BC}}\). Theorem 13.4.55 gives the corresponding factorization of the marginal support inverse. The scalar factors cancel in the Petz sandwich. The identity follows first for \(A_0\otimes X_B\) and then for every operator by a finite sum of product operators.

If \(P_B\) is the support projector of \(\rho _B\), then the support projector of \(d_A^{-1}\mathbf1_A\otimes \rho _B\) is

\begin{align} P_{AB} & =\mathbf1_A\otimes P_B. \label{eq:entropy_mixed_reference_support} \end{align}
Proof

Insert the factorized support inverse into \(P_{AB}=\sigma _{AB,\mathrm{supp}}^{-1/2} \sigma _{AB}\sigma _{AB,\mathrm{supp}}^{-1/2}\). The scalar factors cancel, and the remaining sandwich is \(\mathbf1_A\otimes (\rho _{B,\mathrm{supp}}^{-1/2}\rho _B \rho _{B,\mathrm{supp}}^{-1/2})=\mathbf1_A\otimes P_B\).

Theorem 13.4.58 Complete positivity of the support Petz map

The support Petz transpose map \(\mathcal R_\sigma \) is completely positive.

Proof

For an orthonormal basis \((e_r)_r\) of \(H_R\), let \(J_r:H_L\to H_L\otimes H_R\) be given by \(J_r(v)=v\otimes e_r\). Then

\begin{align} \mathcal R_\sigma (X) & =\sum _r K_rXK_r^\dagger , \notag \\ K_r & =\sqrt\sigma \, J_r\tau ^{-1/2}_{\mathrm{supp}}. \notag \end{align}

Thus the support Petz map has a rectangular Kraus representation.

Lemma 13.4.59 Trace of the support Petz map

Let \(P_\tau \) be the orthogonal projector onto the support of \(\tau =\operatorname{tr}_R\sigma \). Then, for every matrix \(X\),

\begin{align} \operatorname{tr}(\mathcal R_\sigma (X)) & =\operatorname{tr}(P_\tau X). \label{eq:entropy_support_petz_trace} \end{align}
Proof

Cyclicity of the trace and the defining property of the partial trace give

\begin{align} \operatorname{tr}(\mathcal R_\sigma (X)) & =\operatorname{tr}(\tau ^{-1/2}_{\mathrm{supp}} X\tau ^{-1/2}_{\mathrm{supp}}\tau ) =\operatorname{tr}((\tau ^{-1/2}_{\mathrm{supp}} \tau \tau ^{-1/2}_{\mathrm{supp}})X). \notag \end{align}

By (55), this equals \(\operatorname{tr}(P_\tau X)\).

Definition 13.4.60 Complementary preparation term

Put \(Q_\tau =\mathbf1_L-P_\tau \) and \(\omega _R=\operatorname{tr}_L\sigma \). The complementary term is

\begin{align} \mathcal C_\sigma (X) & =Q_\tau XQ_\tau \otimes \omega _R. \label{eq:entropy_petz_complement} \end{align}
Theorem 13.4.61 Complete positivity of the complementary term

The complementary preparation term \(\mathcal C_\sigma \) is completely positive.

Proof

The map \(X\mapsto Q_\tau XQ_\tau \) has the single Kraus operator \(Q_\tau \). Adjoining the positive semidefinite matrix \(\omega _R\) is completely positive, and the composition of these two maps is completely positive.

Lemma 13.4.62 Trace of the complementary term

If \(\sigma \) has trace one, then

\begin{align} \operatorname{tr}(\mathcal C_\sigma (X)) & =\operatorname{tr}(Q_\tau X). \label{eq:entropy_petz_complement_trace} \end{align}
Proof

Since \(\operatorname{tr}(\omega _R)=\operatorname{tr}(\sigma )=1\) and \(Q_\tau ^2=Q_\tau \), factorization and cyclicity of the trace give \(\operatorname{tr}(\mathcal C_\sigma (X)) =\operatorname{tr}(Q_\tau XQ_\tau ) =\operatorname{tr}(Q_\tau ^2X) =\operatorname{tr}(Q_\tau X)\).

Definition 13.4.63 Trace-preserving Petz channel for a partial trace

The completed Petz map is

\begin{align} \widehat{\mathcal R}_\sigma & =\mathcal R_\sigma +\mathcal C_\sigma . \label{eq:entropy_completed_petz} \end{align}

If \(\sigma \) is positive semidefinite with trace one, then \(\widehat{\mathcal R}_\sigma \) is completely positive and trace preserving.

Proof

The two summands are completely positive. Since \(\operatorname{tr}(\omega _R)=1\), the complementary term satisfies (78). Adding (78) to (76) gives trace preservation.

For the reference in (71), the chosen complementary preparation term factors as

\begin{align} \mathcal C_{\sigma _{ABC}} & =\operatorname {id}_A\otimes \mathcal C_{\rho _{BC}}. \label{eq:entropy_petz_complement_factorization} \end{align}

after the canonical reassociation of the three tensor factors. This follows from the chosen support completion in (77), not from [ HJPW04 , equation (10) ] .

Proof

From (75), \(Q_{AB}=\mathbf1_A\otimes Q_B\), while \(\operatorname{tr}_{AB}\sigma _{ABC}=\rho _C\). Hence

\begin{align} \mathcal C_{\sigma _{ABC}}(A_0\otimes X_B) & =(\mathbf1_A\otimes Q_B)(A_0\otimes X_B) (\mathbf1_A\otimes Q_B)\otimes \rho _C \notag \\ & =A_0\otimes \mathcal C_{\rho _{BC}}(X_B). \notag \end{align}

A finite product-operator decomposition gives the result for every input.

For the reference in (71),

\begin{align} \widehat{\mathcal R}_{\sigma _{ABC}} & =\operatorname {id}_A\otimes \widehat{\mathcal R}_{\rho _{BC}}. \label{eq:entropy_completed_petz_factorization} \end{align}

after canonical reassociation of the three tensor factors. The raw support-map summand is the maximally mixed specialization of [ HJPW04 , equation (10) ] . The complementary summand comes from the chosen support completion in (77). This theorem does not assert a Hayashi–Koashi–Imoto decomposition.

Proof

Add (74) and (80), and use additivity of the tensor product of linear maps.

If \(\rho _{BC}\) is positive semidefinite with trace one, then \(\widehat{\mathcal R}_{\sigma _{ABC}}\) is completely positive and trace preserving.

Proof

By (72), \(\operatorname{tr}(\sigma _{ABC})=\operatorname{tr}(\rho _{BC})=1\). Apply the channel property of the completed Petz map.

If \(P_\tau XP_\tau =X\), then \(\widehat{\mathcal R}_\sigma (X)=\mathcal R_\sigma (X)\).

Proof

The assumption implies \(Q_\tau XQ_\tau =0\). Hence \(\mathcal C_\sigma (X)=0\), and adding the complementary term to \(\mathcal R_\sigma (X)\) does not change its value.

Definition 13.4.69 Support projection of a Hermitian matrix

Let \(A\) be Hermitian, with spectral decomposition \(A=U\operatorname{diag}(\lambda _i)U^\dagger \). Its support projection is

\begin{align} P_A & =U\operatorname{diag}(\mathbf1_{\lambda _i\ne 0})U^\dagger . \label{eq:entropy_hermitian_support} \end{align}

The logarithm identity below is a generic support-projection fact, with no tensor-network content; it is relocated here from the MPDO area-law chapter, whose tripartite strong-subadditivity argument cites it across the chapter boundary.

Theorem 13.4.70 Logarithm of a tensor product on its support

Let \(A\) and \(B\) be positive semidefinite, with \(P_A\) and \(P_B\) the orthogonal projections onto their respective ranges. Then

\begin{align} \log (A\otimes B) & =(\log A)\otimes P_B+P_A\otimes (\log B). \label{eq:mpdo_support_log_tensor_product} \end{align}

Here the logarithm is extended by zero on the kernel. Neither matrix is required to be positive definite, and either index set may be empty. This auxiliary identity is project-derived rather than a theorem of CPSV16.

Proof

Diagonalize \(A\) and \(B\). On a product eigenvector with eigenvalues \(a,b\geq 0\), equation (83) becomes

\begin{align} \log (ab) & =\log (a)\mathbf1_{b\ne 0} +\mathbf1_{a\ne 0}\log (b). \notag \end{align}

If \(a,b{\gt}0\), this is the ordinary product identity for the logarithm. If either eigenvalue is zero, both sides vanish. Conjugating the resulting diagonal identity by the product eigenbasis proves the formula.

Theorem 13.4.71 Support projection absorption under kernel inclusion

Let \(A\) and \(B\) be Hermitian matrices such that \(\ker A\subseteq \ker B\). If \(P_A\) is the support projection of \(A\), then \(P_A B P_A=B\).

Proof

The complementary projection \(\mathbf1-P_A\) has range contained in \(\ker A\), and hence in \(\ker B\). Thus \(B(\mathbf1-P_A)=0\). Taking adjoints gives \((\mathbf1-P_A)B=0\), so \(B=P_A B\), and therefore \(P_A B P_A=P_A B=B\).

Definition 13.4.72 Lift along one basis vector
#

For \(w\in H_L\) and a distinguished basis vector \(e_r\in H_R\), define the lift \(J_r w=w\otimes e_r\in H_L\otimes H_R\).

Theorem 13.4.73 Matrix action on a basis-vector lift

Let \(X\) be a matrix on \(H_L\otimes H_R\). For basis indices \(i,s\) and \(r\),

\begin{align} [XJ_r w]_{(i,s)} & =\sum _j X_{(i,s),(j,r)}w_j. \label{eq:entropy_lift_action} \end{align}
Proof

Since \((J_r w)_{(j,c)}=w_j\mathbf1_{c=r}\), expansion of the matrix-vector product gives

\begin{align} [XJ_r w]_{(i,s)} & =\sum _{j,c}X_{(i,s),(j,c)}[J_r w]_{(j,c)} =\sum _j X_{(i,s),(j,r)}w_j. \notag \end{align}
Theorem 13.4.74 Quadratic form of a right marginal

Let \(X\) be a matrix on \(H_L\otimes H_R\), let \(w\in H_L\), and let \((e_r)_r\) be the distinguished orthonormal basis of \(H_R\). Then

\begin{align} \langle w,(\operatorname{tr}_R X)w\rangle & =\sum _r\langle w\otimes e_r,X(w\otimes e_r)\rangle . \label{eq:entropy_partial_trace_qform} \end{align}
Proof

Expanding the matrix products and the partial trace gives \(\sum _{i,j,r}\overline{w_i}X_{(i,r),(j,r)}w_j\) on both sides.

Theorem 13.4.75 Kernel inclusion descends under partial trace

Let \(\sigma \) be positive semidefinite on \(H_L\otimes H_R\), and let \(\rho \) be any matrix on the same space such that \(\ker \sigma \subseteq \ker \rho \). Then

\begin{align} \ker (\operatorname{tr}_R\sigma ) & \subseteq \ker (\operatorname{tr}_R\rho ). \label{eq:entropy_partial_trace_kernel} \end{align}
Proof

If \(w\in \ker (\operatorname{tr}_R\sigma )\), then

\begin{align} 0 & =\langle w,(\operatorname{tr}_R\sigma )w\rangle =\sum _r\langle w\otimes e_r,\sigma (w\otimes e_r)\rangle . \notag \end{align}

Each summand is non-negative and therefore vanishes. Positive semidefiniteness shows that \(\sigma (w\otimes e_r)=0\) for every \(r\), so the joint kernel inclusion gives \(\rho (w\otimes e_r)=0\). Summing the diagonal components over \(r\) yields \((\operatorname{tr}_R\rho )w=0\).

Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(H_L\otimes H_R\) such that \(\ker \sigma \subseteq \ker \rho \). Then

\begin{align} \widehat{\mathcal R}_\sigma (\operatorname{tr}_R\rho ) & =\mathcal R_\sigma (\operatorname{tr}_R\rho ). \label{eq:entropy_petz_marginal} \end{align}
Proof

Kernel inclusion descends through the partial trace, giving \(\ker (\operatorname{tr}_R\sigma )\subseteq \ker (\operatorname{tr}_R\rho )\). Theorem 13.4.71 therefore yields \(P_\tau (\operatorname{tr}_R\rho )P_\tau =\operatorname{tr}_R\rho \), where \(\tau =\operatorname{tr}_R\sigma \). By Theorem 13.4.68,

\begin{align} P_\tau (\operatorname{tr}_R\rho )P_\tau =\operatorname{tr}_R\rho \quad \Longrightarrow \quad \widehat{\mathcal R}_\sigma (\operatorname{tr}_R\rho ) & =\mathcal R_\sigma (\operatorname{tr}_R\rho ). \notag \end{align}
Theorem 13.4.77 The support Petz map recovers its reference

For every positive semidefinite \(\sigma \),

\begin{align} \mathcal R_\sigma (\operatorname{tr}_R\sigma ) & =\sigma . \label{eq:entropy_support_petz_recovery} \end{align}
Proof

By (55), the middle factor reduces to \(P_\tau \otimes \mathbf1_R\). Marginal-support absorption gives \((\mathbf1_{LR}-P_\tau \otimes \mathbf1_R)\sigma =0\), so

\begin{align} \sqrt\sigma (\mathbf1_{LR}-P_\tau \otimes \mathbf1_R)\sqrt\sigma & =0, \notag \end{align}

because that matrix is positive semidefinite of trace zero. Therefore \(\sqrt\sigma (P_\tau \otimes \mathbf1_R)\sqrt\sigma =\sigma \).

Theorem 13.4.78 The completed Petz channel recovers its reference

For every positive semidefinite \(\sigma \),

\begin{align} \widehat{\mathcal R}_\sigma (\operatorname{tr}_R\sigma ) & =\sigma . \label{eq:entropy_completed_petz_recovery} \end{align}
Proof

The marginal satisfies \(P_\tau \tau P_\tau =\tau \). Therefore the completed channel agrees with the support Petz map at \(\tau \), and Theorem 13.4.77 gives \(\widehat{\mathcal R}_\sigma (\tau )=\sigma \).

Theorem 13.4.79 Entropy is invariant under reindexing

Let \(\rho \) be a Hermitian matrix indexed by a finite set \(J\), and let \(e : I \to J\) be a bijection from a finite set \(I\). The reindexed matrix on \(I\) with entries \(\rho _{e(i)\, e(j)}\) has the same von Neumann entropy as \(\rho \).

Proof

By Lemma 13.4.9 the entropy depends only on the characteristic polynomial. Reindexing conjugates \(\rho \) by a permutation matrix, so \(\chi _{(\rho _{e(i)\, e(j)})}=\chi _\rho \), and the entropies coincide.

Lemma 13.4.80 Cyclic invariance of the charpoly-root entropy sum
#

Let \(A \in M_{m \times n}(\mathbb {C})\) and \(B \in M_{n \times m}(\mathbb {C})\). Then the charpoly-root entropy sum is invariant under the cyclic swap \(AB \mapsto BA\):

\begin{align} \sum _{\lambda \in \mathrm{roots}(\chi _{AB})} ({-}\operatorname{Re}(\lambda )\log \operatorname{Re}(\lambda )) & = \sum _{\mu \in \mathrm{roots}(\chi _{BA})} ({-}\operatorname{Re}(\mu )\log \operatorname{Re}(\mu )), \label{eq:entropy_charpoly_cyclic} \end{align}

with roots counted with algebraic multiplicity.

Proof

The rectangular characteristic-polynomial identity gives \(X^n\chi _{AB}=X^m\chi _{BA}\). Thus the two characteristic polynomials have the same nonzero roots, with multiplicity; the additional zero roots contribute nothing because \(0\log 0=0\).

Theorem 13.4.81 Entropy of \(AB\) equals entropy of \(BA\)

For matrices \(A \in M_{m \times n}(\mathbb {C})\) and \(B \in M_{n \times m}(\mathbb {C})\) such that \(AB\) and \(BA\) are Hermitian, \(S(AB)=S(BA)\).

Proof

By Lemma 13.4.9, the entropy of a Hermitian matrix is its charpoly-root entropy sum. Lemma 13.4.80 identifies these two sums for \(AB\) and \(BA\).

Theorem 13.4.82 Entropy is additive over tensor products
#

For density matrices \(\omega \) and \(\tau \), \(S(\omega \otimes \tau )=S(\omega )+S(\tau )\).

Proof

The tensor product is unitarily conjugate to the diagonal of eigenvalue products \(\lambda _i\mu _j\). Using

\begin{align} -\lambda _i\mu _j\log (\lambda _i\mu _j) & =\mu _j\, (-\lambda _i\log \lambda _i) +\lambda _i\, (-\mu _j\log \mu _j) \notag \end{align}

and the unit eigenvalue sums collapses the double sum to \(S(\omega )+S(\tau )\).

Theorem 13.4.83 Entropy of a scaled density matrix
#

For a density matrix \(\omega \) and a scalar \(c\), \(S(c\, \omega )=c\, S(\omega )-c\log c\).

Proof

The eigenvalues of \(c\, \omega \) are \(c\lambda _i\), so the entropy is \(\sum _i-(c\lambda _i)\log (c\lambda _i)\). The splitting identity

\begin{align} -(c\lambda )\log (c\lambda ) & =\lambda \, (-c\log c)+c\, (-\lambda \log \lambda ) \notag \end{align}

together with the unit eigenvalue sum \(\sum _i\lambda _i=1\) gives

\begin{align} \sum _i-(c\lambda _i)\log (c\lambda _i) & =(-c\log c)\sum _i\lambda _i +c\sum _i(-\lambda _i\log \lambda _i) =-c\log c+c\, S(\omega ), \notag \end{align}

the stated form.

Theorem 13.4.84 Entropy is additive over an orthogonal direct sum of Hermitian blocks
#

For a family of Hermitian matrices \(M_j\), the block-diagonal direct sum satisfies

\begin{align} S\! \left(\bigoplus _j M_j\right) & =\sum _j S(M_j). \label{eq:entropy_block_diagonal} \end{align}
Proof

Each block diagonalizes by a unitary congruence, and the block-diagonal assembly of the block unitaries diagonalizes the direct sum, whose eigenvalue multiset is the disjoint union of the block eigenvalue multisets. Hence

\begin{align} S\! \left(\bigoplus _j M_j\right) & =\sum _{\lambda \in \biguplus _j\mathrm{spec}(M_j)} -\lambda \log \lambda =\sum _j\sum _{\lambda \in \mathrm{spec}(M_j)} -\lambda \log \lambda =\sum _j S(M_j). \notag \end{align}
Theorem 13.4.85 Entropy of summands on pairwise-annihilating supports

Let \(A=\sum _j M_j\) be a finite sum of Hermitian matrices. Suppose there are operators \(P_j\) such that

\begin{align} P_jM_j=M_jP_j & =M_j, \notag \\ P_jP_k & =0 \quad (j\ne k). \notag \end{align}

Then \(S(A)=\sum _j S(M_j)\). The operators \(P_j\) need only resolve the support of \(A\); their sum need not be the identity on the ambient space. This is the support form of the direct-sum entropy identity used in [ CPGSV16 , Appendix C.2, lines 1760–1770 ] .

Proof

Write \(A=XY\), where \(X\) maps the direct sum of the labelled ambient spaces to the original space by the matrices \(M_j\), and \(Y\) maps back by the operators \(P_j\). The support and annihilation identities give \(YX=\bigoplus _jM_j\). Reversing the two rectangular factors preserves the nonzero eigenvalues, while any additional eigenvalues are zero. Entropy is therefore unchanged, and additivity on the block diagonal gives the result.

Theorem 13.4.86 Entropy of a weighted orthogonal direct sum

Let \(\omega _j\) be density matrices and let \(p_j\geq 0\). Then

\begin{align} S\! \left(\bigoplus _j p_j\omega _j\right) & =\sum _j(-p_j\log p_j+p_jS(\omega _j)). \label{eq:entropy_weighted_blocks} \end{align}

In particular, when the \(p_j\) form a probability distribution, this is

\begin{align} S\! \left(\bigoplus _j p_j\omega _j\right) & =H(p)+\sum _j p_jS(\omega _j). \label{eq:entropy_weighted_probability} \end{align}

This is the entropy identity used in [ CPGSV16 , Appendix C.2, lines 1760–1770 ] .

Proof

Entropy is additive over the orthogonal blocks. Applying the scaled-state formula to each block gives \(S(p_j\omega _j)=-p_j\log p_j+p_jS(\omega _j)\), and summing over \(j\) gives the result.

Theorem 13.4.87 Equality in a positively weighted sum
#

Suppose \(p_j{\gt}0\) and \(L_j\leq R_j\) for every \(j\). If \(\sum _jp_jL_j=\sum _jp_jR_j\), then \(L_j=R_j\) for every \(j\). This is the positivity argument applied to strong subadditivity in [ CPGSV16 , Appendix C.2, lines 1770–1780 ] .

Proof

Each number \(p_j(R_j-L_j)\) is non-negative, and their sum vanishes. Therefore every one vanishes. Since \(p_j{\gt}0\), it follows that \(R_j-L_j=0\).

13.5 Tripartite partial traces

Definition 13.5.1 Partial trace over \(A\)
#

For a tripartite matrix \(\rho _{ABC}\) on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B} \otimes \mathbb {C}^{d_C}\), the partial trace over \(A\) is

\begin{align} (\operatorname{tr}_A \rho _{ABC})_{(b_1, c_1)(b_2, c_2)} & = \sum _{a=0}^{d_A - 1} (\rho _{ABC})_{(a, b_1, c_1)(a, b_2, c_2)}. \label{eq:entropy_trace_a} \end{align}
Definition 13.5.2 Partial trace over \(C\)
#

The partial trace over \(C\) is

\begin{align} (\operatorname{tr}_C \rho _{ABC})_{(a_1, b_1)(a_2, b_2)} & = \sum _{c=0}^{d_C - 1} (\rho _{ABC})_{(a_1, b_1, c)(a_2, b_2, c)}. \label{eq:entropy_trace_c} \end{align}
Definition 13.5.3 Partial trace over \(AC\)
#

The partial trace over \(A\) and \(C\) is

\begin{align} (\operatorname{tr}_{AC} \rho _{ABC})_{b_1 b_2} & = \sum _{a=0}^{d_A - 1} \sum _{c=0}^{d_C - 1} (\rho _{ABC})_{(a, b_1, c)(a, b_2, c)}. \label{eq:entropy_trace_ac} \end{align}
Lemma 13.5.4 Partial trace preserves Hermiticity

If \(\rho _{ABC}\) is Hermitian, then \(\operatorname{tr}_A(\rho _{ABC})\), \(\operatorname{tr}_C(\rho _{ABC})\), and \(\operatorname{tr}_{AC}(\rho _{ABC})\) are all Hermitian. The same holds for bipartite partial traces \(\operatorname{tr}_A(\rho _{AB})\) and \(\operatorname{tr}_B(\rho _{AB})\).

Proof

Follows from \(\overline{\rho _{ji}} = \rho _{ij}\) applied entry-wise inside the summation defining each partial trace.

13.6 Strong subadditivity

Lemma 13.6.1 Relative entropy against a maximally mixed reference

Let \(\rho \) be a density matrix on \(A \otimes R\) with reduced state \(\rho _R = \operatorname{tr}_A \rho \), and suppose the support condition \(\ker ((\mathbb {1}_A / d_A) \otimes \rho _R) \subseteq \ker \rho \) holds. Then

\begin{align} D(\rho \big\| (\mathbb {1}_A / d_A) \otimes \rho _R) & = \log d_A + S(\rho _R) - S(\rho ). \label{eq:entropy_mixed_reference} \end{align}
Proof

When \(\rho _R\) is singular the tensor logarithm of the reference does not split. Regularize the reduced state by the affine perturbation \(\rho _{R,\varepsilon } = (1 + d_R\varepsilon )^{-1}(\rho _R + \varepsilon \mathbb {1})\), positive definite for \(\varepsilon {\gt} 0\). The reference \((\mathbb {1}_A / d_A) \otimes \rho _{R,\varepsilon }\) is then a positive definite tensor product, whose logarithm splits, so the cross trace term evaluates to \(-\log d_A + \operatorname{Re}\operatorname{tr}(\rho _R\log \rho _{R,\varepsilon })\). As \(\varepsilon \to 0^+\) the support condition makes the zero eigenvalues of \(\rho _R\) contribute nothing, so

\begin{align} \operatorname{Re}\operatorname{tr}(\rho \, \log ((\mathbb {1}_A / d_A) \otimes \rho _{R,\varepsilon })) & \to -\log d_A - S(\rho _R). \notag \end{align}

Since \(D(\rho \| \sigma ) = -S(\rho ) - \operatorname{Re}\operatorname{tr}(\rho \log \sigma )\), the evaluation follows.

Lemma 13.6.2 The maximally mixed reference lies in the support domain
#

Let \(\rho \) be a positive semidefinite operator on \(A \otimes R\) with reduced state \(\rho _R = \operatorname{tr}_A \rho \). Then the singular reference \((\mathbb {1}_A / d_A) \otimes \rho _R\) satisfies the support condition \(\ker ((\mathbb {1}_A / d_A) \otimes \rho _R) \subseteq \ker \rho \).

Proof

A vector \(v\) annihilated by \((\mathbb {1}_A / d_A) \otimes \rho _R\) is annihilated by \(\mathbb {1}_A \otimes \rho _R\), since \(\mathbb {1}_A / d_A\) is invertible. Let \(P\) be the orthogonal projection onto the range of \(\rho _R\). The complementary lift \(\mathbb {1}_A \otimes (\mathbb {1}- P)\) then fixes \(v\), while it annihilates \(\rho \) on the left because the reduced state of \(\rho \) on \(R\) is supported on the range of \(P\). Hence \(\rho v = 0\).

Lemma 13.6.3 Partial trace of the maximally mixed reference

For every tripartite matrix \(\rho _{ABC}\),

\begin{align} \operatorname{tr}_C\! \left(\frac{\mathbb {1}_A}{d_A}\otimes \rho _{BC}\right) & =\frac{\mathbb {1}_A}{d_A}\otimes \rho _B. \label{eq:entropy_trace_mixed} \end{align}
Proof

For indices \((a_1,b_1)\) and \((a_2,b_2)\), the corresponding matrix entry is

\begin{align} \left[\operatorname{tr}_C\! \left(\frac{\mathbb {1}_A}{d_A}\otimes \rho _{BC}\right)\right]_{(a_1,b_1),(a_2,b_2)} & =\frac{\delta _{a_1,a_2}}{d_A} \sum _c (\rho _{BC})_{(b_1,c),(b_2,c)} =\frac{\delta _{a_1,a_2}}{d_A}(\rho _B)_{b_1,b_2}. \notag \end{align}

This is the corresponding entry of \((\mathbb {1}_A/d_A)\otimes \rho _B\).

For a tripartite density matrix \(\rho _{ABC}\),

\begin{align} S(\rho _{ABC}) + S(\rho _B) & \le S(\rho _{AB}) + S(\rho _{BC}). \label{eq:entropy_ssa} \end{align}
Proof

Read the inequality as one instance of data processing under the partial trace over \(C\), with the singular reference state \(\sigma _{ABC} = (\mathbb {1}_A / d_A) \otimes \rho _{BC}\). Being a density operator is not enough to place the pair \((\rho _{ABC}, \sigma _{ABC})\) in the relative-entropy domain, which is the kernel inclusion \(\ker \sigma _{ABC} \subseteq \ker \rho _{ABC}\); the marginal support lemma supplies it. Against this reference the relative entropy of each pair evaluates to an entropy difference:

\begin{align} D(\rho _{ABC} \big\| (\mathbb {1}_A / d_A) \otimes \rho _{BC}) & = \log d_A + S(\rho _{BC}) - S(\rho _{ABC}), \label{eq:entropy_ssa_eval_abc}\\ D(\rho _{AB} \big\| (\mathbb {1}_A / d_A) \otimes \rho _B) & = \log d_A + S(\rho _B) - S(\rho _{AB}). \label{eq:entropy_ssa_eval_ab} \end{align}

For a singular reference the tensor logarithm does not split, so each evaluation regularizes the reduced state through the affine perturbation \(\rho _{R,\varepsilon } = (1 + d_R\varepsilon )^{-1}(\rho _R + \varepsilon \mathbb {1})\), which is positive definite, and passes to the limit

\begin{align} \operatorname{Re}\operatorname{tr}(\rho \, \log ((\mathbb {1}_A / d_A) \otimes \rho _{R,\varepsilon })) & \to -\log d_A - S(\rho _R) \qquad (\varepsilon \to 0^+), \notag \end{align}

where the zero eigenvalues of \(\rho _R\) contribute nothing because the kernel inclusion makes the corresponding diagonal weights vanish. Data processing on the singular support domain under the partial trace over \(C\), which sends \(\rho _{ABC} \mapsto \rho _{AB}\) and \((\mathbb {1}_A / d_A) \otimes \rho _{BC} \mapsto (\mathbb {1}_A / d_A) \otimes \rho _B\), reads

\begin{align} D(\rho _{AB} \big\| (\mathbb {1}_A / d_A) \otimes \rho _B) & \le D\bigl(\rho _{ABC} \big\| (\mathbb {1}_A / d_A) \otimes \rho _{BC}\bigr). \label{eq:entropy_ssa_dpi} \end{align}

Substituting (100) and (101) into (102) cancels the common \(\log d_A\) and rearranges to the claimed inequality.

For a positive definite tripartite density matrix \(\rho _{ABC}\),

\begin{align} S(\rho _{ABC}) + S(\rho _B) & \le S(\rho _{AB}) + S(\rho _{BC}). \notag \end{align}

This is the positive definite case of Theorem 13.6.4.

Proof

Read the inequality as one instance of data processing under the partial trace over \(C\), with reference state \(\sigma _{ABC} = (\mathbb {1}_A / d_A) \otimes \rho _{BC}\), which is positive definite because \(\rho _{BC}\) is. Against this reference the relative entropy of each pair evaluates to an entropy difference:

\begin{align} D(\rho _{ABC} \big\| (\mathbb {1}_A / d_A) \otimes \rho _{BC}) & = \log d_A + S(\rho _{BC}) - S(\rho _{ABC}), \label{eq:entropy_ssa_posdef_abc}\\ D(\rho _{AB} \big\| (\mathbb {1}_A / d_A) \otimes \rho _B) & = \log d_A + S(\rho _B) - S(\rho _{AB}). \label{eq:entropy_ssa_posdef_ab} \end{align}

Each equality follows from the tensor logarithm split and the adjoint of the partial trace. Data processing under the partial trace over \(C\), which sends \(\rho _{ABC} \mapsto \rho _{AB}\) and \((\mathbb {1}_A / d_A) \otimes \rho _{BC} \mapsto (\mathbb {1}_A / d_A) \otimes \rho _B\), reads

\begin{align} D(\rho _{AB} \big\| (\mathbb {1}_A / d_A) \otimes \rho _B) & \le D\bigl(\rho _{ABC} \big\| (\mathbb {1}_A / d_A) \otimes \rho _{BC}\bigr). \label{eq:entropy_ssa_posdef_dpi} \end{align}

Substituting (103) and (104) into (105) cancels the common \(\log d_A\) and rearranges to the claimed inequality.

Definition 13.6.6 SSA equality

A tripartite density matrix \(\rho _{ABC}\) satisfies SSA equality if

\begin{align} S(\rho _{ABC}) + S(\rho _B) & = S(\rho _{AB}) + S(\rho _{BC}). \label{eq:entropy_ssa_equality} \end{align}

Let \(\rho \) and \(\sigma \) be positive semidefinite matrices such that \(\ker \sigma \subseteq \ker \rho \), and let \(\tau \) be positive semidefinite with \(\operatorname{tr}\tau =1\). Then

\begin{align} D(\rho \otimes \tau \, \Vert \, \sigma \otimes \tau ) & =D(\rho \, \Vert \, \sigma ). \label{eq:entropy_ancilla_additivity_support} \end{align}
Proof

Write \(P_\rho \), \(P_\sigma \), and \(P_\tau \) for the support projections. The kernel inclusion gives \(P_\sigma \rho P_\sigma =\rho \) and \(\rho P_\sigma =\rho \), while \(\rho P_\rho =\rho \) and \(\tau P_\tau =\tau \). Substituting the singular tensor-logarithm formulas into the relative entropy and using these four identities cancels the two contributions containing \(\log \tau \). Factoring the trace of the remaining tensor product gives

\begin{align} D(\rho \otimes \tau \, \Vert \, \sigma \otimes \tau ) & =D(\rho \, \Vert \, \sigma )\operatorname{tr}\tau =D(\rho \, \Vert \, \sigma ). \notag \end{align}

Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\) such that \(\ker \sigma \subseteq \ker \rho \), and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For a primitive \(d_C\)-th root of unity, put \(U_{ab}=\mathbb {1}_S\otimes W(a,b)\) and

\begin{align} \overline X & =\frac{1}{d_C^2}\sum _{c,e}U_{ce}XU_{ce}^{\dagger }. \label{eq:entropy_weyl_average} \end{align}

Then, for every \(a,b\),

\begin{align} D(\overline\rho \Vert \overline\sigma ) & =D(U_{ab}\rho U_{ab}^{\dagger }\Vert U_{ab}\sigma U_{ab}^{\dagger }). \label{eq:entropy_weyl_summand} \end{align}

This is a scalar equality-propagation prerequisite for [ HJPW04 , Theorem 3 and equation (8) ] ; it neither characterizes equality in joint convexity nor asserts recovery.

Proof

Let \(\tau _C=d_C^{-1}\mathbb {1}_C\). The twirl identity gives \(\overline X=(\operatorname{tr}_C X)\otimes \tau _C\). The support inclusion passes to the partial traces:

\begin{align} \ker \sigma \subseteq \ker \rho & \Longrightarrow \ker (\operatorname{tr}_C\sigma )\subseteq \ker (\operatorname{tr}_C\rho ). \notag \end{align}

Hence support-domain ancilla additivity and the saturation hypothesis give

\begin{align} D(\overline\rho \Vert \overline\sigma ) & =D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma ) =D(\rho \Vert \sigma ). \label{eq:entropy_weyl_average_value} \end{align}

Every \(U_{ab}\) is unitary, and unitary invariance gives

\begin{align} D(U_{ab}\rho U_{ab}^{\dagger }\Vert U_{ab}\sigma U_{ab}^{\dagger }) & =D(\rho \Vert \sigma ). \label{eq:entropy_weyl_conjugate} \end{align}

Together, (110) and (111) prove (109).

Let \(\rho \) and \(\sigma \) be positive semidefinite matrices on \(\mathcal{H}_S\otimes \mathbb {C}^{d_C}\) such that \(\ker \sigma \subseteq \ker \rho \), and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For a primitive \(d_C\)-th root of unity, put \(U_{ce}=\mathbb {1}_S\otimes W(c,e)\) and

\begin{align} \overline X & =\frac{1}{d_C^2}\sum _{c,e}U_{ce}XU_{ce}^{\dagger }. \notag \end{align}

Then

\begin{align} D(\overline\rho \Vert \overline\sigma ) & =\frac{1}{d_C^2}\sum _{c,e} D(U_{ce}\rho U_{ce}^{\dagger }\Vert U_{ce}\sigma U_{ce}^{\dagger }). \label{eq:entropy_weyl_jensen} \end{align}

This scalar identity is associated with the finite Jensen step in the Weyl proof of data processing. It is a prerequisite for [ HJPW04 , Theorem 3 and equation (8) ] ; it neither characterizes equality in joint convexity nor asserts recovery.

Proof

Unitary invariance gives, for every \(c,e\),

\begin{align} D(U_{ce}\rho U_{ce}^{\dagger }\Vert U_{ce}\sigma U_{ce}^{\dagger }) & =D(\rho \Vert \sigma ). \notag \end{align}

Theorem 13.6.8 gives \(D(\overline\rho \Vert \overline\sigma )=D(\rho \Vert \sigma )\). Substituting into the right-hand side gives

\begin{align} \frac{1}{d_C^2}\sum _{c,e} D(U_{ce}\rho U_{ce}^{\dagger }\Vert U_{ce}\sigma U_{ce}^{\dagger }) & =\frac{d_C^2}{d_C^2}D(\rho \Vert \sigma ) =D(\overline\rho \Vert \overline\sigma ). \notag \end{align}
Lemma 13.6.10 Positive homogeneity of relative entropy

If \(A\) and \(B\) are positive definite and \(c{\gt}0\), then

\begin{align} D(cA\Vert cB) & =cD(A\Vert B). \label{eq:entropy_homogeneity} \end{align}
Proof

The functional-calculus identity \(\log (cA)=(\log c)\mathbb {1}+\log A\), and its analog for \(B\), show that the scalar logarithmic terms cancel in \(\log (cA)-\log (cB)\). Linearity of the trace then gives the result.

Lemma 13.6.11 Support-domain positive homogeneity of relative entropy

Let \(A\) and \(B\) be positive semidefinite and suppose that \(\ker B\subseteq \ker A\). If \(c{\gt}0\), then

\begin{align} D(cA\Vert cB) & =cD(A\Vert B). \label{eq:entropy_support_homogeneity} \end{align}
Proof

On the support of a positive semidefinite matrix,

\begin{align} \log (cA)& =(\log c)P_A+\log A. \notag \end{align}

The support projection \(P_A\) fixes \(A\) on the right. The kernel inclusion also makes \(P_B\) fix \(A\) on the right. Hence the two terms containing \(\log c\) cancel in the trace-log difference, and linearity of the trace gives the result.

Theorem 13.6.12 Partial-trace saturation gives zero weighted Weyl gap

Let \(\rho \) and \(\sigma \) be positive definite, and suppose that \(D(\rho \Vert \sigma )=D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). For the uniformly weighted Weyl conjugates \(A_g=d_C^{-2}U_g\rho U_g^\dagger \) and \(B_g=d_C^{-2}U_g\sigma U_g^\dagger \), put \(A=\sum _gA_g\) and \(B=\sum _gB_g\). Then

\begin{align} \sum _gD(A_g\Vert B_g)-D(A\Vert B) & =0. \label{eq:entropy_weyl_gap} \end{align}
Proof

Positive definiteness of \(\sigma \) makes the support condition in Theorem 13.6.9 automatic. That theorem gives the equality with \(d_C^{-2}\) multiplying each scalar relative entropy. Lemma 13.6.10 moves this coefficient inside both matrix arguments, which is precisely the stated weighted gap.

Lemma 13.6.13 Scalar logarithmic resolvent integral

Let \(a,b{\gt}0\). Then the function

\begin{align} t & \longmapsto \frac{a^2/(a+tb)-b+t b^2/(a+tb)}{1+t} \notag \end{align}

is integrable on \((0,\infty )\), and

\begin{align} \int _0^\infty \frac{a^2/(a+tb)-b+t b^2/(a+tb)}{1+t}\, dt & =a(\log a-\log b). \label{eq:entropy_scalar_resolvent} \end{align}

This is the scalar normalization \((\mathrm{intspec})\) in Jenčová–Ruskai, arXiv:0903.2895v4, §2.1, lines 406–413.

Proof

The integrand is the derivative of \(a(\log (1+t)-\log (a+tb))\). Its value at \(t=0\) is \(-a\log a\), whereas its limit as \(t\to \infty \) is \(-a\log b\). The derivative has constant sign, according as \(a-b\) is positive or negative, and is therefore integrable. The fundamental theorem of calculus gives the stated value.

Theorem 13.6.14 Hermitian trace-log spectral identity

Let \(A\) and \(B\) be Hermitian matrices of the same size, with spectral resolutions

\begin{align} A& =\sum _i\alpha _i|u_i\rangle \! \langle u_i|, \notag \\ B& =\sum _j\beta _j|v_j\rangle \! \langle v_j|. \notag \end{align}

Write \(w_{ij}=\langle u_i,v_j\rangle \), and use the total real logarithm: \(\log 0=0\), while \(\log x=\log |x|\) for \(x{\lt}0\). Then

\begin{align} D(A\Vert B) & =\sum _{i,j}\alpha _i (\log \alpha _i-\log \beta _j)|w_{ij}|^2. \label{eq:entropy_hermitian_spectral} \end{align}

This is an algebraic totalized extension to arbitrary Hermitian matrices of the homogeneous trace-log identity \((\mathrm{J1})\), which Jenčová–Ruskai state for strictly positive matrices in arXiv:0903.2895v4, lines 277–287.

Proof

Expand both trace terms in eigenbases. Unitarity of the overlap matrix gives \(\sum _j|w_{ij}|^2=1\), so

\begin{align} \operatorname{Re}\operatorname{tr}(A\log A) & =\sum _{i,j}\alpha _i\log (\alpha _i)|w_{ij}|^2, \notag \\ \operatorname{Re}\operatorname{tr}(A\log B) & =\sum _{i,j}\alpha _i\log (\beta _j)|w_{ij}|^2. \notag \end{align}

Subtraction gives (117).

Theorem 13.6.15 Support-domain spectral integral of relative entropy

Let \(A\) and \(B\) be positive semidefinite matrices of the same size and suppose that \(\ker B\subseteq \ker A\). With the spectral notation of Theorem 13.6.14, define, for \(t{\gt}0\),

\begin{align} r_{ij}(t) & = \begin{cases} 0, & \beta _j=0,\\ \frac{ \alpha _i^2/(\alpha _i+t\beta _j)-\beta _j +t\beta _j^2/(\alpha _i+t\beta _j)}{1+t}, & \beta _j{\gt}0, \end{cases} \notag \\ I_{A,B}(t) & =\sum _{i,j}r_{ij}(t)|w_{ij}|^2. \label{eq:entropy_support_integrand} \end{align}

Thus no ordinary quotient with \(\alpha _i=\beta _j=0\) is used. Define also

\begin{align} e_{ij} & = \begin{cases} 0, & \beta _j=0,\\ \alpha _i(\log \alpha _i-\log \beta _j)|w_{ij}|^2, & \beta _j{\gt}0. \end{cases} \notag \end{align}

The function \(I_{A,B}\) is integrable on \((0,\infty )\), and

\begin{align} \int _0^\infty I_{A,B}(t)\, dt & =\sum _{i,j}e_{ij} =D(A\Vert B). \label{eq:entropy_support_integral} \end{align}

The trace-log identity \((\mathrm{J1})\), the scalar normalization \((\mathrm{intspec})\), and its matrix form \((\mathrm{intAB})\) occur in Jenčová–Ruskai, arXiv:0903.2895v4, at lines 277–287, 406–413, and 423–427, respectively. The support-domain extension is given at lines 717–720.

This theorem concerns the spectral expression \(I_{A,B}\). The next theorem identifies it with the coordinate-free left-right quadratic form underlying the finite Weyl formula.

Proof

If \(\beta _j=0\), the kernel inclusion gives \(\alpha _i|w_{ij}|^2=0\), and both \(r_{ij}\) and \(e_{ij}\) are defined to be zero. If \(\beta _j{\gt}0\), the scalar integral applies when \(\alpha _i{\gt}0\), while the term is identically zero when \(\alpha _i=0\). Since the double sum is finite, integration term by term gives

\begin{align} \int _0^\infty I_{A,B}(t)\, dt & =\sum _{i,j}e_{ij}. \notag \end{align}

The kernel inclusion also shows that replacing each \(\beta _j=0\) summand in Theorem 13.6.14 by zero does not change its value. Hence that theorem identifies the sum in (123) with \(D(A\Vert B)\).

Let \(A\) and \(B\) be positive semidefinite matrices of the same size, let \(P_B\) be the orthogonal projection onto the support of \(B\), and, for \(t{\gt}0\), set

\begin{align} S_t & =A\otimes \mathbb {1}+t(\mathbb {1}\otimes B^{\top }). \notag \end{align}

Write \(S_t^+\) for the generalized inverse that vanishes on \(\ker S_t\). Define

\begin{align} Q_{A P_B}(t) & =\operatorname{Re}\left\langle \operatorname{vec}((A P_B)^{\top }), S_t^+\operatorname{vec}((A P_B)^{\top }) \right\rangle , \notag \\ Q_B(t) & =\operatorname{Re}\left\langle \operatorname{vec}(B^{\top }), S_t^+\operatorname{vec}(B^{\top }) \right\rangle . \notag \end{align}

With the spectral notation of Theorem 13.6.15,

\begin{align} Q_{A P_B}(t) & =\sum _{i,j:\, \beta _j{\gt}0} \frac{\alpha _i^2}{\alpha _i+t\beta _j}|w_{ij}|^2, \label{eq:entropy_support_quad_a}\\ Q_B(t) & =\sum _{i,j:\, \beta _j{\gt}0} \frac{\beta _j^2}{\alpha _i+t\beta _j}|w_{ij}|^2. \label{eq:entropy_support_quad_b} \end{align}

Consequently,

\begin{align} \frac{Q_{A P_B}(t)-\operatorname{Re}\operatorname{tr}B+tQ_B(t)}{1+t} & =I_{A,B}(t). \label{eq:entropy_support_quadratic} \end{align}

The function in (126) is continuous on \((0,\infty )\). If \(\ker B\subseteq \ker A\), then \(A P_B=A\), so \(Q_{A P_B}(t)\) equals the quadratic form with source \(\operatorname{vec}(A^{\top })\).

This is the support-projected form of \((\mathrm{intAB})\) in Jenčová–Ruskai, arXiv:0903.2895v4, §2.1, lines 423–427, with the support convention at lines 717–720.

Proof

Diagonalize \(A\) and \(B\) and write \(W=U_A^\ast U_B\). In these coordinates, the two equations

\begin{align} A X_A+tX_A B & =A P_B, \label{eq:entropy_support_solution_a}\\ A X_B+tX_B B & =B \label{eq:entropy_support_solution_b} \end{align}

have entries

\begin{align} (U_A^\ast X_AU_B)_{ij} & = \begin{cases} \dfrac {\alpha _i}{\alpha _i+t\beta _j}w_{ij}, & \beta _j{\gt}0,\\ 0, & \beta _j=0, \end{cases} \label{eq:entropy_support_entries_a}\\ (U_A^\ast X_BU_B)_{ij} & = \begin{cases} \dfrac {\beta _j}{\alpha _i+t\beta _j}w_{ij}, & \beta _j{\gt}0,\\ 0, & \beta _j=0. \end{cases} \label{eq:entropy_support_entries_b} \end{align}

For every positive semidefinite \(S\), the identities \(S^+S=SS^+=P_S\) imply that \(Sx=b\) gives \(\langle b,S^+b\rangle =\langle b,x\rangle \). Applying this identity to (127) and (128), with (131) and (134), gives (124) and (125). Their coefficientwise combination is the scalar function appearing in Lemma 13.6.13. Continuity follows term by term from the finite sums.

Lemma 13.6.17 Nonnegativity of the finite-family source-\(B\) support defect

Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size. For \(t{\gt}0\), write

\begin{align} Q^B_{A_i,B_i}(t) & = \operatorname{Re}\left\langle \operatorname{vec}(B_i^{\mathsf T}), \bigl(A_i\otimes \mathbf1+ t(\mathbf1\otimes B_i^{\mathsf T})\bigr)^+ \operatorname{vec}(B_i^{\mathsf T}) \right\rangle . \notag \end{align}

Then

\begin{align} 0 & \leq \sum _{i\in I}Q^B_{A_i,B_i}(t) -Q^B_{\sum _i A_i,\sum _i B_i}(t). \notag \end{align}

No kernel inclusion between \(A_i\) and \(B_i\) is required. This lemma is a positive-semidefinite support-domain extension of the positive-definite calculation in equations \((\mathrm{Mj})\), \((\mathrm{eq:Schz1})\), and \((\mathrm{eq:Schwzt})\) at lines 1313–1343 of Jenčová–Ruskai, arXiv:0903.2895v4; their generalized-inverse notation is given at lines 254–262. The paper does not state this extension. Its later singular equality theorem at lines 761–785 assumes \(\ker B_i\subseteq \ker A_i\) and is not asserted here. The subsequent singular entropy-equality passage is recorded in the TNLean paper-gap note [ con26p ] .

Proof

Put

\begin{align} S_i & =A_i\otimes \mathbf1+t(\mathbf1\otimes B_i^{\mathsf T}),& b_i& =\operatorname{vec}(B_i^{\mathsf T}). \notag \end{align}

The source equation for the support relative-modular operator shows that \(b_i\) lies in the support of \(S_i\). The same argument applied to \(\sum _i A_i\) and \(\sum _i B_i\) shows that \(\sum _i b_i\) lies in the support of \(\sum _i S_i\). The support-resolvent residual identity writes the displayed defect as a sum of nonnegative quadratic residuals.

Theorem 13.6.18 Vanishing source-\(B\) defect gives common left–right solutions

Under the hypotheses and notation of Lemma 13.6.17, suppose that the source-\(B\) defect vanishes:

\begin{align} \sum _{i\in I}Q^B_{A_i,B_i}(t) -Q^B_{A_\Sigma ,B_\Sigma }(t) & =0. \notag \end{align}

Put

\begin{align} S_i& =A_i\otimes \mathbf1+ t(\mathbf1\otimes B_i^{\mathsf T}),& S_\Sigma & =\sum _{i\in I}S_i, \notag \end{align}

and let \(P_{S_i}\) be the support projection of \(S_i\). Then

\begin{align} S_i^+\operatorname{vec}(B_i^{\mathsf T}) & = P_{S_i}S_\Sigma ^+\operatorname{vec}(B_\Sigma ^{\mathsf T}) \qquad (i\in I). \notag \end{align}

This is the fixed-parameter support-domain residual step behind \((\mathrm{basiceq})\) in Jenčová–Ruskai, arXiv:0903.2895v4, lines 652–660 and 788–790. No kernel inclusion is needed for this algebraic implication.

Proof

The source vectors lie in the supports of their left–right operators, and their sum lies in the support of \(S_\Sigma \). The vanishing real defect and the support-resolvent residual identity force every quadratic residual to vanish. The common-solution conclusion follows after projection to the support of each \(S_i\).

Lemma 13.6.19 Nonnegativity of the finite-family unprojected source-\(A\) support defect

Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size satisfying \(\ker B_i\subseteq \ker A_i\). For \(t{\gt}0\), write

\begin{align} Q^A_{A_i,B_i}(t) & = \operatorname{Re}\left\langle \operatorname{vec}(A_i^{\mathsf T}), \bigl(A_i\otimes \mathbf1+ t(\mathbf1\otimes B_i^{\mathsf T})\bigr)^+ \operatorname{vec}(A_i^{\mathsf T}) \right\rangle . \notag \end{align}

Then

\begin{align} 0 & \leq \sum _{i\in I}Q^A_{A_i,B_i}(t) -Q^A_{\sum _i A_i,\sum _i B_i}(t). \notag \end{align}

The kernel inclusions for the summands imply \(\ker (\sum _i B_i)\subseteq \ker (\sum _i A_i)\). This lemma is the source-\(A\) support-domain extension of the positive-definite residual calculation in equations \((\mathrm{Mj})\), \((\mathrm{eq:Schz1})\), and \((\mathrm{eq:Schwzt})\) at lines 1313–1343 of Jenčová–Ruskai, arXiv:0903.2895v4. The paper does not state this fixed-parameter extension separately. The lemma does not assert that equality of relative entropies makes the defect vanish.

Proof

Put

\begin{align} S_i & =A_i\otimes \mathbf1+t(\mathbf1\otimes B_i^{\mathsf T}),& a_i& =\operatorname{vec}(A_i^{\mathsf T}). \notag \end{align}

The inclusion \(\ker B_i\subseteq \ker A_i\) gives \(A_iP_{B_i}=A_i\). The source-\(A\) left–right equation therefore places \(a_i\) in the support of \(S_i\). Positivity shows that a vector annihilated by \(\sum _i B_i\) is annihilated by every \(B_i\), and hence by every \(A_i\). Thus the summed source also lies in the support of the summed operator. The support-resolvent residual identity writes the displayed difference as a sum of nonnegative quadratic residuals.

Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices of the same size satisfying \(\ker B_i\subseteq \ker A_i\). Put

\begin{align} A_\Sigma & =\sum _{i\in I}A_i,& B_\Sigma & =\sum _{i\in I}B_i, \notag \\ \operatorname {Def}_A(t) & = \sum _{i\in I}Q^A_{A_i,B_i}(t) -Q^A_{A_\Sigma ,B_\Sigma }(t),& \operatorname {Def}_B(t) & = \sum _{i\in I}Q^B_{A_i,B_i}(t) -Q^B_{A_\Sigma ,B_\Sigma }(t). \notag \end{align}

Then the function

\begin{align} F(t) & = \frac{\operatorname {Def}_A(t)+ t\operatorname {Def}_B(t)}{1+t} \notag \end{align}

is continuous on \((0,\infty )\) and integrable there, and

\begin{align} \int _0^\infty F(t)\, dt & = \sum _{i\in I}D(A_i\Vert B_i) -D(A_\Sigma \Vert B_\Sigma ). \label{eq:support_relative_entropy_family_gap_integral} \end{align}

This is the finite-family support-domain form of \((\mathrm{intspec})\) and \((\mathrm{intAB})\) in Jenčová–Ruskai, arXiv:0903.2895v4, lines 406–431, with the positive-semidefinite convention and kernel hypotheses at lines 717–720 and 766–785.

Proof

The kernel inclusions for the summands imply \(\ker B_\Sigma \subseteq \ker A_\Sigma \). Apply the one-pair support-domain integral representation to every \((A_i,B_i)\) and to \((A_\Sigma ,B_\Sigma )\). The pointwise left–right identities identify the difference of the spectral integrands with \(F\); linearity of trace cancels the terms involving \(\operatorname {tr}(B_i)\). Finite summation commutes with the integral. Continuity follows from the corresponding one-pair continuity theorem and finite summation.

Theorem 13.6.21 Nonnegativity of the finite-family support gap integrand

Under the hypotheses and notation of Theorem 13.6.20, for every \(t{\gt}0\) one has

\begin{align} 0 & \leq \frac{\operatorname {Def}_A(t)+ t\operatorname {Def}_B(t)}{1+t}. \notag \end{align}

This is the coefficient-correct combination of the two source defects in \((\mathrm{intAB})\) of Jenčová–Ruskai, arXiv:0903.2895v4, lines 423–435.

Proof

Both defects are nonnegative. Since \(t{\gt}0\) and \(1+t{\gt}0\), multiplying the source-\(B\) defect by \(t\), adding the source-\(A\) defect, and dividing by \(1+t\) preserves nonnegativity.

Theorem 13.6.22 Relative-entropy equality annihilates the source-\(B\) defect

Under the hypotheses of Theorem 13.6.20, suppose

\begin{align} D(A_\Sigma \Vert B_\Sigma ) & = \sum _{i\in I}D(A_i\Vert B_i). \notag \end{align}

Then, for every \(t{\gt}0\),

\begin{align} \operatorname {Def}_B(t) & = \sum _{i\in I}Q^B_{A_i,B_i}(t) -Q^B_{A_\Sigma ,B_\Sigma }(t) =0. \notag \end{align}

This is the source-\(B\) pointwise-vanishing passage in Jenčová–Ruskai, arXiv:0903.2895v4, lines 433–435, 652–674, and 788–790. The common projected resolvent conclusion is stated downstream in Theorem 13.6.35.

Proof

The relative-entropy equality and (135) make the integral of the nonnegative function \(F\) vanish. Hence \(F=0\) almost everywhere. Its continuity on the open positive half-line upgrades this to \(F(t)=0\) for every \(t{\gt}0\). Both defects are nonnegative, so \(\operatorname {Def}_A(t)+t\operatorname {Def}_B(t)=0\), and \(t{\gt}0\) forces \(\operatorname {Def}_B(t)=0\).

Let \(A\) and \(B\) be positive definite, with spectral resolutions \(A=\sum _i\alpha _i|u_i\rangle \! \langle u_i|\) and \(B=\sum _j\beta _j|v_j\rangle \! \langle v_j|\). Put \(w_{ij}=\langle u_i,v_j\rangle \). Let \(L_A\) and \(R_B\) denote left and right multiplication, \(L_A(X)=AX\) and \(R_B(X)=XB\). Then, for \(t{\gt}0\),

\begin{align} \operatorname{Re}\langle A,(L_A+tR_B)^{-1}A\rangle _{\rm HS} & =\sum _{i,j}\frac{\alpha _i^2}{\alpha _i+t\beta _j}|w_{ij}|^2, \label{eq:entropy_resolvent_a}\\ \operatorname{Re}\langle B,(L_A+tR_B)^{-1}B\rangle _{\rm HS} & =\sum _{i,j}\frac{\beta _j^2}{\alpha _i+t\beta _j}|w_{ij}|^2. \label{eq:entropy_resolvent_b} \end{align}

Both quadratic forms are continuous on \((0,\infty )\). Moreover,

\begin{align} D(A\Vert B) & =\int _0^\infty \frac{ \operatorname{Re}\langle A,(L_A+tR_B)^{-1}A\rangle _{\rm HS} -\operatorname{Re}\operatorname{tr}B +t\operatorname{Re}\langle B,(L_A+tR_B)^{-1}B\rangle _{\rm HS}}{1+t}\, dt, \label{eq:entropy_resolvent_integral} \end{align}

and the integrand in (138) is continuous and integrable on the positive half-line. This is the positive-definite spectral route of Jenčová–Ruskai, arXiv:0903.2895v4, §4.

Proof

Vectorization sends \(L_A+tR_B\) to \(A\otimes \mathbb {1}+t\mathbb {1}\otimes B^{\mathsf T}\). The vectors \(u_i\otimes \overline{v_j}\) diagonalize this matrix with eigenvalues \(\alpha _i+t\beta _j\), which proves (136) and (137). Insert these identities into (138), use \(\sum _i|w_{ij}|^2=1\), and apply Lemma 13.6.13 term by term. The same finite spectral sum proves continuity and integrability.

Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(S_i\) be a positive-semidefinite matrix, let \(P_i\) be its support projection, and let \(b_i\) lie in its support. Write

\begin{align} S& =\sum _i S_i, \notag \\ b& =\sum _i b_i, \notag \\ G_i& =\left((S_i)^{-1/2}_{\mathrm{supp}}\right)^2, \notag \\ G& =\left(S^{-1/2}_{\mathrm{supp}}\right)^2. \notag \end{align}

Assume also that \(b\) lies in the support of \(S\), and put \(x=Gb\). Then

\begin{align} \sum _i \langle b_i-S_i x,G_i(b_i-S_i x)\rangle & =\sum _i\langle b_i,G_i b_i\rangle -\langle b,Gb\rangle . \label{eq:entropy_residual_identity} \end{align}

In particular,

\begin{align} 0 & \leq \operatorname{Re}\left( \sum _i\langle b_i,G_i b_i\rangle -\langle b,Gb\rangle \right). \notag \end{align}

Jenčová and Ruskai give the positive-definite residual expansion in equations \((\mathrm{Mj})\) and \((\mathrm{eq:Schz1})\) of the Appendix to arXiv:0903.2895v4. The support assumptions make the same expansion valid for the generalized inverses of the \(S_i\) and of \(S\).

Proof

Since \(G_iS_i=S_iG_i=P_i\) and \(P_i b_i=b_i\), expansion of the \(i\)th summand gives

\begin{align} \langle b_i,G_i b_i\rangle -\langle b_i,x\rangle -\langle x,b_i\rangle +\langle x,S_i x\rangle . \notag \end{align}

Sum over \(i\). The support assumption on \(b\) gives \(Sx=SGb=b\), so the last three terms combine to \(-\langle b,Gb\rangle \).

Theorem 13.6.25 Vanishing real support-resolvent defect gives a common solution

Under the hypotheses and notation of Lemma 13.6.24, suppose that

\begin{align} \operatorname{Re}\left( \sum _i\langle b_i,G_i b_i\rangle -\langle b,Gb\rangle \right) & =0. \label{eq:entropy_zero_defect} \end{align}

Then, for every \(i\in I\),

\begin{align} G_i b_i & =P_i x. \label{eq:entropy_common_solution} \end{align}

This is the support-domain form of the common-resolvent equation \((\mathrm{basiceq})\) in Section 3.1 of Jenčová–Ruskai, arXiv:0903.2895v4. Its residual calculation is the one in the Appendix, equations \((\mathrm{Mj})\) and \((\mathrm{eq:Schz1})\).

Proof

Lemma 13.6.24 writes the complex defect in (140) as a finite sum of non-negative quadratic forms. Its imaginary part therefore vanishes automatically, while the hypothesis makes its real part vanish. Hence \(G_i(b_i-S_ix)=0\) for every \(i\). Multiplication by \(S_i\) shows that \(P_i(b_i-S_ix)=0\). Both \(b_i\) and \(S_ix\) lie in the support of \(S_i\), so \(b_i=S_ix\). Multiplication by \(G_i\) now gives (141).

Theorem 13.6.26 Fixed-resolvent defect identity for the Weyl family

Let \(\rho \) and \(\sigma \) be positive definite on \(\mathcal H_S\otimes \mathbb C^{d_C}\), let \(q=d_C^{-2}\), and put

\begin{align} A_g& =qU_g\rho U_g^{\dagger },& B_g& =qU_g\sigma U_g^{\dagger },& A& =\sum _g A_g,& B& =\sum _g B_g, \notag \end{align}

where \(U_g=\mathbf1_S\otimes W_g\). For \(t{\gt}0\), let

\begin{align} T_g(X)& =A_gX+tXB_g,& T(X)& =AX+tXB,& \Lambda _t& =T^{-1}(B). \notag \end{align}

Then

\begin{align} \sum _g\left\| T_g^{-1/2}(B_g)-T_g^{1/2}(\Lambda _t) \right\| _{\mathrm{HS}}^2 & = \sum _g\operatorname{Re}\langle B_g,T_g^{-1}(B_g)\rangle _{\mathrm{HS}} -\operatorname{Re}\langle B,T^{-1}(B)\rangle _{\mathrm{HS}}. \notag \end{align}

This is the positive-definite, fixed-\(t\) identity in Jenčová–Ruskai, arXiv:0903.2895v4, Appendix, lines 1313–1343. It does not infer zero defect from equality of relative entropies and makes no assertion about singular supports.

Proof

Write \(S_g\) for the positive definite matrix representing \(T_g\) under the vectorization \(X\mapsto \operatorname{vec}(X^{\mathsf T})\), and put \(b_g=\operatorname{vec}(B_g^{\mathsf T})\) and \(x=S^{-1}\sum _g b_g\). For the residual \(r_g=b_g-S_gx\), direct expansion gives

\begin{align} \sum _g\langle r_g,S_g^{-1}r_g\rangle & = \sum _g\langle b_g,S_g^{-1}b_g\rangle -\left\langle \sum _gb_g, S^{-1}\sum _gb_g\right\rangle . \notag \end{align}

Since \(S_g^{1/2}\) is invertible,

\begin{align} \| S_g^{-1/2}b_g-S_g^{1/2}x\| ^2 & = \| S_g^{-1/2}r_g\| ^2 =\operatorname{Re}\langle r_g,S_g^{-1}r_g\rangle . \notag \end{align}

Summing proves the identity.

Theorem 13.6.27 Source-\(A\) fixed-resolvent defect identity

Under the notation of Theorem 13.6.26, set \(\Gamma _t=T^{-1}(A)\). Then, for every \(t{\gt}0\),

\begin{align} \sum _g\left\| T_g^{-1/2}(A_g)-T_g^{1/2}(\Gamma _t) \right\| _{\mathrm{HS}}^2 & = \sum _g\operatorname{Re}\langle A_g,T_g^{-1}(A_g)\rangle _{\mathrm{HS}} -\operatorname{Re}\langle A,T^{-1}(A)\rangle _{\mathrm{HS}}. \notag \end{align}

This is the second defect family in Jenčová–Ruskai, arXiv:0903.2895v4, §4 and Appendix.

Proof

Apply the residual identity of Theorem 13.6.26 with \(a_g=\operatorname{vec}(A_g^{\mathsf T})\) and \(\bar a=\operatorname{vec}(A^{\mathsf T})\). Thus

\begin{align} \sum _g\langle a_g,S_g^{-1}a_g\rangle -\langle \bar a,S^{-1}\bar a\rangle & = \sum _g\| S_g^{-1/2}a_g-S_g^{1/2}S^{-1}\bar a\| ^2, \notag \end{align}

which is the asserted source-\(A\) identity.

For the positive definite finite-Weyl family, let

\begin{align} \operatorname{Def}_A(t) & = \sum _g\operatorname{Re}\langle A_g,T_g^{-1}(A_g)\rangle _{\mathrm{HS}} -\operatorname{Re}\langle A,T^{-1}(A)\rangle _{\mathrm{HS}}, \notag \end{align}

and define \(\operatorname{Def}_B(t)\) analogously, with the real parts of the corresponding source-\(B\) pairings. Then both defects are non-negative and continuous for \(t{\gt}0\), and

\begin{align} \sum _gD(A_g\Vert B_g)-D(A\Vert B) & = \int _0^\infty \frac{\operatorname{Def}_A(t)+t\operatorname{Def}_B(t)}{1+t}\, dt. \notag \end{align}

The integrand is continuous, non-negative, and integrable on \((0,\infty )\). The coefficient of the source-\(B\) defect is exactly \(t\), as prescribed by the integral formula and equality analysis of Jenčová–Ruskai, arXiv:0903.2895v4, §4.

Proof

Apply Theorem 13.6.23 to each pair \((A_g,B_g)\) and to \((A,B)\), and interchange the finite sum with the integral. The trace terms cancel because \(B=\sum _gB_g\). The two fixed-resolvent identities write the defects as sums of squared norms, proving nonnegativity. Continuity and integrability follow from the corresponding spectral assertions before taking the finite difference.

Theorem 13.6.29 Zero Weyl defect gives the common resolvent solution

Under the hypotheses and notation of the preceding theorem, suppose that

\begin{align} \sum _g\operatorname{Re}\langle B_g,T_g^{-1}(B_g)\rangle _{\mathrm{HS}} -\operatorname{Re}\langle B,T^{-1}(B)\rangle _{\mathrm{HS}} & =0. \label{eq:entropy_weyl_zero_defect} \end{align}

Then \(T_g^{-1}(B_g)=T^{-1}(B)\) for every Weyl index \(g\). In particular this holds for the identity Weyl element \(g=(0,0)\). This is the common-resolvent conclusion in Jenčová–Ruskai, arXiv:0903.2895v4, §4, lines 652–674; its squared-defect input is in Appendix, lines 1313–1343. This fixed-\(t\), positive-definite conclusion does not assert that equality of relative entropies implies the scalar hypothesis in (142).

Proof

The hypothesis (142) is a finite sum of squared norms. Each term is non-negative, so every term vanishes. Thus \(S_g^{-1/2}(b_g-S_gx)=0\) for every \(g\). Invertibility of \(S_g^{-1/2}\) gives \(b_g=S_gx\), and hence \(S_g^{-1}b_g=x=S^{-1}\sum _hb_h\).

Under the positive-definite finite-Weyl hypotheses, suppose that \(\sum _gD(A_g\Vert B_g)-D(A\Vert B)=0\). Then \(\operatorname{Def}_B(t)=0\) for every \(t{\gt}0\), and consequently \(T_g^{-1}(B_g)=T^{-1}(B)\) for every \(t{\gt}0\) and every Weyl index \(g\). In particular, equality of relative entropy under the right partial trace implies this conclusion by Theorem 13.6.12. This is the positive-definite conclusion of the equality argument in Jenčová–Ruskai, arXiv:0903.2895v4, §4 and Appendix. It makes no assertion at \(t=0\) or for singular inputs.

Proof

The integrand in Theorem 13.6.28 is non-negative and has integral zero, hence it vanishes almost everywhere. Its continuity improves this to vanishing at every \(t{\gt}0\). Since \(\operatorname{Def}_A(t)\geq 0\), \(\operatorname{Def}_B(t)\geq 0\), and \(t/(1+t){\gt}0\), it follows that \(\operatorname{Def}_B(t)=0\). Theorem 13.6.29 now gives the common solution.

Let \(S,T\) be positive semidefinite matrices and let \(x\) be a vector. If \((t\mathbf1+S)^{-1}x=(t\mathbf1+T)^{-1}x\) for every \(t{\gt}0\), then \(\sqrt S\, x=\sqrt T\, x\). More generally, for any fixed matrix \(Q\), if \(Q(t\mathbf1+S)^{-1}x=Q(t\mathbf1+T)^{-1}x\) for every \(t{\gt}0\), then \(Q\sqrt S\, x=Q\sqrt T\, x\).

In particular, for positive definite \(A,B\) and every \(t{\gt}0\), put \(\Delta _{A,B}=A\otimes (B^{-1})^{\mathsf T}\). The source-\(B\) left–right resolvent satisfies

\begin{align} \left(A\otimes \mathbf1 +t\, \mathbf1\otimes B^{\mathsf T}\right)^{-1} \operatorname{vec}(B^{\mathsf T}) & = \left(t\mathbf1+\Delta _{A,B}\right)^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}), \notag \end{align}

and

\begin{align} \sqrt{\Delta _{A,B}} \operatorname{vec}(\mathbf1^{\mathsf T}) & = \operatorname{vec}\! \left((\sqrt A\, (\sqrt B)^{-1})^{\mathsf T}\right). \notag \end{align}

These are the positive-square-root specializations of the passage from relative modular resolvents to analytic functions of the relative modular operator in Jenčová–Ruskai, arXiv:0903.2895v4, lines 658–680.

Proof

Use the Löwner integral representation of the power \(p=1/2\). Its integrand at \(t{\gt}0\) is \(f_t(S)=t^{-1/2}\mathbf1-t^{1/2}(t\mathbf1+S)^{-1}\). Applied to \(x\), this is \(f_t(S)x=t^{-1/2}x-t^{1/2}(t\mathbf1+S)^{-1}x\). The hypothesis therefore gives \(Q(f_t(S)x)=Q(f_t(T)x)\) for every \(t{\gt}0\). Since \(M\mapsto Q(Mx)\) is a bounded linear map from \(M_{n}(\mathbb {C})\) to \(\mathbb C^n\),

\begin{align} Q\left(\left(\int _0^\infty f_t(S)\, d\mu (t)\right)x\right) & = \int _0^\infty Q(f_t(S)x)\, d\mu (t), \notag \end{align}

and similarly for \(T\). Integration therefore gives \(Q\sqrt S\, x=Q\sqrt T\, x\).

For the second assertion, factor \(L_A+tR_B=(\Delta _{A,B}+t\mathbf1)R_B\). Applying the inverse to \(B=R_B(\mathbf1)\) gives the shifted relative modular resolvent on \(\mathbf1\). Finally,

\begin{align} \sqrt{\Delta _{A,B}} & = \sqrt A\otimes ((\sqrt B)^{-1})^{\mathsf T}, \notag \end{align}

as follows from uniqueness of the positive square root.

Lemma 13.6.32 Right multiplication under transposed vectorization

For matrices \(M\) and \(P\), column-stacking vectorization satisfies

\begin{align} \left(\mathbf1\otimes P^{\mathsf T}\right) \operatorname{vec}(M^{\mathsf T}) & = \operatorname{vec}\! \left((MP)^{\mathsf T}\right). \notag \end{align}
Proof

This is the Kronecker vectorization identity \((B\otimes A)\operatorname{vec}(X)=\operatorname{vec}(AXB^{\mathsf T})\) with \(A=P^{\mathsf T}\), \(B=\mathbf1\), and \(X=M^{\mathsf T}\).

Lemma 13.6.33 Canonical support-domain source resolvent solution

Let \(A,B\) be positive semidefinite, let \(t{\gt}0\), set \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\), and let \(P_B\) be the support projection of \(B\). Then

\begin{align} & \left(A\otimes \mathbf1 +t(\mathbf1\otimes B^{\mathsf T})\right) (\mathbf1\otimes P_B^{\mathsf T}) \left(t\mathbf1+A\otimes (B^+)^{\mathsf T}\right)^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}) \notag \\ & \hspace{4em}=\operatorname{vec}(B^{\mathsf T}). \notag \end{align}
Proof

Put \(C=\mathbf1\otimes B^{\mathsf T}\). The support generalized-inverse identity gives

\begin{align} C\left(A\otimes (B^+)^{\mathsf T}\right) & =A\otimes P_B^{\mathsf T}. \notag \end{align}

Together with \(B^{\mathsf T}P_B^{\mathsf T}=B^{\mathsf T}\), this factors the left two operators as

\begin{align} \left(A\otimes \mathbf1+tC\right) (\mathbf1\otimes P_B^{\mathsf T}) & = C\left(t\mathbf1+A\otimes (B^+)^{\mathsf T}\right). \notag \end{align}

The shifted relative-modular matrix is positive definite and hence invertible, so

\begin{align} \left(A\otimes \mathbf1+tC\right) (\mathbf1\otimes P_B^{\mathsf T}) \left(t\mathbf1+A\otimes (B^+)^{\mathsf T}\right)^{-1} & =C. \notag \end{align}

The Kronecker vectorization identity gives \(C\operatorname{vec}(\mathbf1^{\mathsf T})=\operatorname{vec}(B^{\mathsf T})\), which is the claim.

Let \(A,B\) be positive semidefinite, let \(t{\gt}0\), set

\begin{align} S& =A\otimes \mathbf1+t(\mathbf1\otimes B^{\mathsf T}),& R& =t\mathbf1+A\otimes (B^+)^{\mathsf T}, \notag \end{align}

and let \(P_B\) be the support projection of \(B\). Then

\begin{align} S^+\operatorname{vec}(B^{\mathsf T}) & = (\mathbf1\otimes P_B^{\mathsf T})R^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}). \notag \end{align}

Consequently,

\begin{align} \operatorname{Re}\left\langle \operatorname{vec}(B^{\mathsf T}),S^+\operatorname{vec}(B^{\mathsf T}) \right\rangle & = \operatorname{Re}\left\langle \operatorname{vec}(B^{\mathsf T}), (\mathbf1\otimes P_B^{\mathsf T})R^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}) \right\rangle . \notag \end{align}

This is the one-pair algebraic identification used in the singular equality argument of Jenčová–Ruskai, arXiv:0903.2895v4, lines 783–790. It does not assert that equality of relative entropies gives a common resolvent for a finite family.

Proof

Put

\begin{align} P& =\mathbf1\otimes P_B^{\mathsf T},& C& =\mathbf1\otimes B^{\mathsf T},& D& =\mathbf1\otimes (B^+)^{\mathsf T}. \notag \end{align}

The generalized-inverse identities give \(SP=CR\), \(CD=P\), and \(P^2=P\). The preceding lemma shows that \(y=PR^{-1}\operatorname{vec}(\mathbf1^{\mathsf T})\) satisfies \(Sy=\operatorname{vec}(B^{\mathsf T})\), and \(Py=y\). Every vector \(v\) satisfying \(Pv=v\) lies in the range of \(S\), since

\begin{align} S\bigl(PR^{-1}Dv\bigr)& =CDv=Pv=v. \notag \end{align}

In particular, \(y\) lies in the range of \(S\), so the support projection \(P_S\) of \(S\) satisfies \(P_Sy=y\). Therefore

\begin{align} S^+\operatorname{vec}(B^{\mathsf T})=S^+Sy=P_Sy=y. \notag \end{align}

Pairing this vector equality with \(\operatorname{vec}(B^{\mathsf T})\) and taking real parts gives the quadratic identity.

Theorem 13.6.35 Relative-entropy equality gives common projected shifted relative-modular resolvents

Let \(I\) be a finite nonempty set. For each \(i\in I\), let \(A_i\) and \(B_i\) be positive-semidefinite matrices satisfying \(\ker B_i\subseteq \ker A_i\). Put

\begin{align} A_\Sigma & =\sum _{i\in I}A_i,& B_\Sigma & =\sum _{i\in I}B_i, \notag \end{align}

and suppose that

\begin{align} D(A_\Sigma \Vert B_\Sigma ) & = \sum _{i\in I}D(A_i\Vert B_i). \notag \end{align}

Then, for every \(i\in I\) and \(t{\gt}0\),

\begin{align} & (\mathbf1\otimes P_{B_i}^{\mathsf T}) \bigl(t\mathbf1+ A_i\otimes (B_i^+)^{\mathsf T}\bigr)^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}) \notag \\ & \qquad = (\mathbf1\otimes P_{B_i}^{\mathsf T}) \bigl(t\mathbf1+ A_\Sigma \otimes (B_\Sigma ^+)^{\mathsf T}\bigr)^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}). \notag \end{align}

Thus the shifted relative-modular resolvents agree on \((\ker B_i)^\perp \). The local support projection is essential; no ambient equality is asserted. This is the support-restricted conclusion of Jenčová–Ruskai, arXiv:0903.2895v4, lines 766–793.

Proof

Relative-entropy equality makes the source-\(B\) defect vanish for every positive parameter. The common left–right solution theorem and the one-pair projected relative-modular identity then compare the local solution with the summed solution. The local right-support projection absorbs both the support projection of the local left–right operator and the right-support projection of the summed reference matrix. Applying it to the common-solution identity gives the displayed equality.

Let \(\rho \) and \(\sigma \) be positive semidefinite matrices satisfying \(\ker \sigma \subseteq \ker \rho \), and suppose that

\begin{align} D(\rho \Vert \sigma ) & =D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma ). \notag \end{align}

Fix a primitive \(d_C\)-th root of unity, put \(U_g=\mathbf1\otimes W_g\), and define the unweighted Weyl family

\begin{align} A_g& =U_g\rho U_g^\dagger ,& B_g& =U_g\sigma U_g^\dagger ,& A_\Sigma & =\sum _g A_g,& B_\Sigma & =\sum _g B_g. \notag \end{align}

Then, for every Weyl index \(g\) and every \(t{\gt}0\),

\begin{align} & (\mathbf1\otimes P_{B_g}^{\mathsf T}) \bigl(t\mathbf1+ A_g\otimes (B_g^+)^{\mathsf T}\bigr)^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}) \notag \\ & \qquad = (\mathbf1\otimes P_{B_g}^{\mathsf T}) \bigl(t\mathbf1+ A_\Sigma \otimes (B_\Sigma ^+)^{\mathsf T}\bigr)^{-1} \operatorname{vec}(\mathbf1^{\mathsf T}). \notag \end{align}

In particular, the zero Weyl index gives the projected resolvent of the original pair \((\rho ,\sigma )\). No ambient resolvent equality is asserted. The support convention is that of Jenčová–Ruskai, arXiv:0903.2895v4, lines 717–720, and their common-resolvent equality argument is at lines 766–793. This is one analytic step toward the recovery implication of Hayden–Jozsa–Petz–Winter, Theorem 3 and equation (8); it is not the direct-sum Markov decomposition invoked in CPSV16, Lemma Lsigma3.

Proof

Unitary conjugation preserves positive semidefiniteness and transports the kernel inclusion to every pair \((A_g,B_g)\). The kernel of the sum of the \(B_g\) is contained in the kernel of the sum of the \(A_g\). The finite Weyl Jensen equality and Lemma 13.6.11 give

\begin{align} D(A_\Sigma \Vert B_\Sigma ) & =\sum _g D(A_g\Vert B_g). \notag \end{align}

Theorem 13.6.35 applied to this finite family gives the displayed projected resolvent.

Lemma 13.6.37 Unweighted finite-Weyl sum normalization

Let \(X\) be a matrix on \(\mathcal H_S\otimes \mathbb C^{d_C}\) and let \(U_g=\mathbf1\otimes W_g\) for the \(d_C^2\) Weyl indices. Then

\begin{align} \sum _g U_g XU_g^\dagger & = d_C^2\bigl((\operatorname{tr}_CX)\otimes d_C^{-1}\mathbf1_C\bigr). \notag \end{align}
Proof

The uniform Weyl twirl is \((\operatorname{tr}_CX)\otimes d_C^{-1}\mathbf1_C\). Multiplying its defining identity by \(d_C^2\) gives the unweighted sum.

Lemma 13.6.38 Support relative-modular square-root evaluation

Let \(A\) and \(B\) be positive semidefinite, and write \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\). Then

\begin{align} \sqrt{A\otimes (B^+)^{\mathsf T}} \operatorname{vec}(\mathbf1^{\mathsf T}) & = \operatorname{vec}\! \left((\sqrt A\, B^{-1/2}_{\mathrm{supp}})^{\mathsf T}\right). \notag \end{align}
Proof

The support inverse square root is positive semidefinite, and

\begin{align} \left(\sqrt A\otimes (B^{-1/2}_{\mathrm{supp}})^{\mathsf T}\right)^2 & = A\otimes (B^+)^{\mathsf T}. \notag \end{align}

Uniqueness of the positive square root gives the corresponding Kronecker factorization. Applying the Kronecker vectorization identity to \(\operatorname{vec}(\mathbf1^{\mathsf T})\) gives the stated equation.

Let \(A,B,C,D\) be positive semidefinite matrices. Write \(B^+=(B^{-1/2}_{\mathrm{supp}})^2\) and \(D^+=(D^{-1/2}_{\mathrm{supp}})^2\), and let \(P_B\) be the orthogonal projection onto \((\ker B)^\perp \). If, for every \(t{\gt}0\),

\begin{align} \left[\left(t\mathbf1+A\otimes (B^+)^{\mathsf T}\right)^{-1} (\mathbf1)\right]P_B & = \left[\left(t\mathbf1+C\otimes (D^+)^{\mathsf T}\right)^{-1} (\mathbf1)\right]P_B, \notag \end{align}

where the inverses act as relative-modular superoperators on matrices, then

\begin{align} \left(\sqrt A\, B^{-1/2}_{\mathrm{supp}}\right)P_B & = \left(\sqrt C\, D^{-1/2}_{\mathrm{supp}}\right)P_B. \notag \end{align}

This is the square-root specialization of the support functional-calculus passage in Jenčová–Ruskai, arXiv:0903.2895v4, lines 788–793. The projection \(P_B\) is essential: the source gives the common generalized resolvents only after restriction to \((\ker B)^\perp \). Deriving this restricted equality requires the preceding singular equality argument and its kernel hypotheses; these are supplied for finite families by Theorem 13.6.35.

Proof

By Lemma 13.6.38,

\begin{align} \sqrt{A\otimes (B^+)^{\mathsf T}} \operatorname{vec}(\mathbf1^{\mathsf T}) & = \operatorname{vec}\! \left((\sqrt A\, B^{-1/2}_{\mathrm{supp}})^{\mathsf T}\right). \notag \end{align}

Setting \(Q=\mathbf1\otimes P_B^{\mathsf T}\) and applying Lemma 13.6.31 gives

\begin{align} \operatorname{vec}\! \left((\sqrt A\, B^{-1/2}_{\mathrm{supp}}P_B)^{\mathsf T}\right) & = \operatorname{vec}\! \left((\sqrt C\, D^{-1/2}_{\mathrm{supp}}P_B)^{\mathsf T}\right). \notag \end{align}

Since \(\operatorname{vec}(X^{\mathsf T})=\operatorname{vec}(Y^{\mathsf T})\) implies \(X=Y\), injectivity of vectorization gives the conclusion.

Theorem 13.6.40 Positive-definite saturation gives the identity Weyl sandwich

Let \(\rho \) and \(\sigma \) be positive definite and suppose that \(D(\rho \Vert \sigma ) =D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma )\). Put

\begin{align} \overline\rho & = (\operatorname{tr}_C\rho )\otimes d_C^{-1}\mathbf1_C, & \overline\sigma & = (\operatorname{tr}_C\sigma )\otimes d_C^{-1}\mathbf1_C. \notag \end{align}

Then

\begin{align} \sqrt\rho \, \sigma ^{-1/2} & = \sqrt{\overline\rho } \overline\sigma ^{-1/2}, \notag \\ \sqrt\sigma \, \overline\sigma ^{-1/2} \overline\rho \, \overline\sigma ^{-1/2} \sqrt\sigma & =\rho . \notag \end{align}

Hence the raw partial-trace Petz map satisfies \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \). This is the positive-definite case only; no singular-support conclusion is asserted.

Proof

Apply the common-resolvent theorem to the identity Weyl summand and the Weyl average. Lemma 13.6.31 converts the result to equality of the two square-root ratios. For \(c=d_C^{-2}\), the uniform scalar in the identity summand cancels through

\begin{align} \sqrt{c\rho }\, (\sqrt{c\sigma })^{-1} & = (\sqrt c\, \sqrt\rho ) ((\sqrt c)^{-1}(\sqrt\sigma )^{-1}) = \sqrt\rho \, (\sqrt\sigma )^{-1}. \notag \end{align}

The Weyl twirl identifies the average with the displayed maximally mixed extensions.

Taking the adjoint product of the ratio equality yields

\begin{align} \sigma ^{-1/2}\rho \sigma ^{-1/2} & = \overline\sigma ^{-1/2}\overline\rho \overline\sigma ^{-1/2}. \notag \end{align}

Multiplication by \(\sqrt\sigma \) on both sides gives the sandwich. The raw Petz recovery identity follows from Theorem 13.6.42.

Lemma 13.6.41 Maximally mixed support-sandwich normalization

Let \(\tau \) be positive semidefinite on \(\mathcal H_S\), let \(X\) be a matrix on \(\mathcal H_S\), and let \(\overline\tau =\tau \otimes d_C^{-1}\mathbf1_C\). Then

\begin{align} \overline\tau ^{-1/2}_{\mathrm{supp}} (X\otimes d_C^{-1}\mathbf1_C) \overline\tau ^{-1/2}_{\mathrm{supp}} & = \bigl(\tau ^{-1/2}_{\mathrm{supp}}X \tau ^{-1/2}_{\mathrm{supp}}\bigr)\otimes \mathbf1_C. \notag \end{align}
Proof

By Lemma 13.4.44, each outer factor contributes \(\sqrt{d_C}\). The scalar coefficient cancels because \(\sqrt{d_C}\, d_C^{-1}\sqrt{d_C} =d_C^{-1}(\sqrt{d_C})^2=1\). The remaining matrix product is the asserted unital tensor embedding.

Theorem 13.6.42 Identity Weyl sandwich implies raw Petz recovery

Let \(\rho \) and \(\sigma \) be matrices on \(\mathcal H_S\otimes \mathbb C^{d_C}\), with \(\sigma \) positive semidefinite, and set \(\overline\sigma =(\operatorname{tr}_C\sigma )\otimes d_C^{-1}\mathbf1_C\). If the identity summand obeys the support sandwich identity

\begin{align} \sqrt\sigma \overline\sigma ^{-1/2}_{\mathrm{supp}} ((\operatorname{tr}_C\rho )\otimes d_C^{-1}\mathbf1_C) \overline\sigma ^{-1/2}_{\mathrm{supp}} \sqrt\sigma & =\rho , \notag \end{align}

then the raw support Petz map recovers \(\rho \): \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \). This is the algebraic reduction in [ HJPW04 , Theorem 3, equation (8) ] . It does not derive the support sandwich identity from equality of relative entropies.

Proof

Lemma 13.6.41 rewrites the middle three factors as

\begin{align} \bigl((\operatorname{tr}_C\sigma )^{-1/2}_{\mathrm{supp}} (\operatorname{tr}_C\rho ) (\operatorname{tr}_C\sigma )^{-1/2}_{\mathrm{supp}}\bigr) \otimes \mathbf1_C. \notag \end{align}

The support Petz formula therefore identifies the left-hand side of the assumed sandwich identity with \(\mathcal R_\sigma (\operatorname{tr}_C\rho )\).

In the finite Weyl coordinates \(\mathbb C^{d_S}\otimes \mathbb C^{\mathbb Z/d_C\mathbb Z}\), let \(\rho \) and \(\sigma \) be positive semidefinite matrices satisfying \(\ker \sigma \subseteq \ker \rho \), and suppose that

\begin{align} D(\rho \Vert \sigma ) & =D(\operatorname{tr}_C\rho \Vert \operatorname{tr}_C\sigma ). \notag \end{align}

Put

\begin{align} \overline\rho & =(\operatorname{tr}_C\rho )\otimes d_C^{-1}\mathbf1_C, & \overline\sigma & =(\operatorname{tr}_C\sigma )\otimes d_C^{-1}\mathbf1_C. \notag \end{align}

If \(P_\sigma \) is the orthogonal projection onto \((\ker \sigma )^\perp \), then

\begin{align} \bigl(\sqrt\rho \, \sigma ^{-1/2}_{\mathrm{supp}}\bigr)P_\sigma & = \bigl(\sqrt{\overline\rho }\, \overline\sigma ^{-1/2}_{\mathrm{supp}}\bigr)P_\sigma , \notag \\ \sqrt\sigma \, \overline\sigma ^{-1/2}_{\mathrm{supp}} \overline\rho \, \overline\sigma ^{-1/2}_{\mathrm{supp}} \sqrt\sigma & =\rho . \notag \end{align}

Consequently, \(\mathcal R_\sigma (\operatorname{tr}_C\rho )=\rho \) for the raw support Petz map. The projected relative-modular argument follows Jenčová–Ruskai, arXiv:0903.2895v4, lines 766–793, and the recovery formula is [ HJPW04 , Theorem 3, equation (8) ] . This result does not assert the middle-space direct-sum decomposition in [ CPGSV16 , Lemma Lsigma3, lines 1351–1363 ] .

Proof

Use the zero Weyl index in Theorem 13.6.36. Lemma 13.6.39 gives the projected square-root ratio for \((\rho ,\sigma )\) and the unweighted Weyl sums. Lemma 13.6.37 identifies those sums with \(d_C^2\overline\rho \) and \(d_C^2\overline\sigma \). The factors \(\sqrt{d_C^2}\) and \((\sqrt{d_C^2})^{-1}\) cancel, giving the first displayed equality.

Taking the adjoint product of that equality gives

\begin{align} P_\sigma \sigma ^{-1/2}_{\mathrm{supp}}\rho \sigma ^{-1/2}_{\mathrm{supp}}P_\sigma & = P_\sigma \overline\sigma ^{-1/2}_{\mathrm{supp}} \overline\rho \, \overline\sigma ^{-1/2}_{\mathrm{supp}}P_\sigma . \notag \end{align}

The support projection absorbs \(\sqrt\sigma \) on both sides, while \(\ker \sigma \subseteq \ker \rho \) gives \(P_\sigma \rho P_\sigma =\rho \). Multiplying the last equality by \(\sqrt\sigma \) on the left and right therefore yields the support sandwich. Theorem 13.6.42 gives the raw Petz recovery identity.

Lemma 13.6.44 Product-coordinate covariance of the partial trace and Petz map

Let \(e_L:H_L\to H_L'\) and \(e_R:H_R\to H_R'\) be bijections, and write \(Z^{e_L\otimes e_R}\) for the corresponding simultaneous relabelling of the rows and columns of \(Z\). Then

\begin{align} \operatorname{tr}_{R'}\bigl(Z^{e_L\otimes e_R}\bigr) & =(\operatorname{tr}_R Z)^{e_L}, \label{eq:petz_product_reindexing_trace}\\ \mathcal R_{\sigma ^{e_L\otimes e_R}} \bigl(X^{e_L}\bigr) & =\bigl(\mathcal R_\sigma (X)\bigr)^{e_L\otimes e_R}. \label{eq:petz_product_reindexing_map} \end{align}
Proof

The first identity follows by changing the summation index in the partial trace. Functional calculus commutes with a relabelling by a bijection, so both \(\sqrt{\sigma }\) and the support inverse square root of \(\operatorname{tr}_R\sigma \) have the same covariance. The embedding \(X\mapsto X\otimes \mathbf1_R\) and matrix multiplication also commute with the product relabelling, which gives (144).

Let \(\rho \) and \(\sigma \) be density operators on the finite-dimensional product \(H_L\otimes H_R\) such that \(\ker \sigma \subseteq \ker \rho \). If

\begin{align} D(\rho \Vert \sigma ) & =D(\operatorname{tr}_R\rho \Vert \operatorname{tr}_R\sigma ), \label{eq:petz_dpi_equality} \end{align}

then the completed partial-trace Petz channel associated with \(\sigma \) recovers \(\rho \):

\begin{align} \widehat{\mathcal R}_\sigma (\operatorname{tr}_R\rho ) & =\rho . \label{eq:petz_dpi_recovery} \end{align}

This is the right-partial-trace forward implication of [ HJPW04 , Theorem 3, equation (8) ] . The additional off-support term belongs to the trace-preserving completion and vanishes on \(\operatorname{tr}_R\rho \); it is not part of the formula in the cited equation.

Proof

Choose bijections from \(H_L\) to a finite standard basis and from \(H_R\) to a finite cyclic basis. Relative entropy, the support inclusion, and the equality (145) are unchanged under this product relabelling. Theorem  13.6.43 gives raw Petz recovery in the cyclic coordinates, and Lemma  13.6.44 transports the identity back to \(H_L\otimes H_R\). Finally, Theorem  13.4.76 identifies the completed channel with the raw support formula at \(\operatorname{tr}_R\rho \), giving (146).

For every positive semidefinite operator \(\omega _{XY}\),

\begin{align} D(\omega _{XY} \| \omega _X\otimes \omega _Y) & = S(\omega _X)+S(\omega _Y)-S(\omega _{XY}). \notag \end{align}

No invertibility assumption is made on either marginal. This is [ HJPW04 , Equation (4) ] .

Proof

The lifted support projections \(P_X\otimes \mathbf1_Y\) and \(\mathbf1_X\otimes P_Y\) both fix \(\omega _{XY}\). Substituting Theorem 13.4.70 into the cross term of the relative entropy and using the two partial-trace adjoint identities gives

\begin{align} \operatorname{Re}\operatorname{tr}(\omega _{XY}\log (\omega _X\otimes \omega _Y)) & = -S(\omega _X)-S(\omega _Y). \notag \end{align}

The asserted formula follows from \(D(\rho \, \| \sigma )=-S(\rho )-\operatorname{Re}\operatorname{tr}(\rho \log \sigma )\).

Lemma 13.6.47 Support of a bipartite state in the product of its marginals

Let \(\omega _{XY}\) be positive semidefinite. Then

\begin{align} \ker (\omega _X\otimes \omega _Y) & \subseteq \ker \omega _{XY}. \notag \end{align}

Equivalently, \(\omega _{XY}\) is supported on \(\operatorname {supp}\omega _X\otimes \operatorname {supp}\omega _Y\).

Proof

The support projection of \(\omega _X\otimes \omega _Y\) is \(P_X\otimes P_Y\). The lifted projections \(P_X\otimes \mathbf1_Y\) and \(\mathbf1_X\otimes P_Y\) both fix \(\omega _{XY}\), so their product \(P_X\otimes P_Y\) fixes it as well. Every vector annihilated by the product of the marginals is therefore annihilated by \(\omega _{XY}\).

Let \(\rho _{ABC}\) be a tripartite density matrix. Then equality holds in strong subadditivity if and only if relative-entropy data processing under the partial trace over \(C\) is saturated for the pair \(\rho _{ABC}\) and \(\rho _A\otimes \rho _{BC}\):

\begin{align} D(\rho _{ABC} \| \rho _A\otimes \rho _{BC}) & = D(\rho _{AB} \| \rho _A\otimes \rho _B). \notag \end{align}

Moreover, \(\operatorname{tr}_C(\rho _A\otimes \rho _{BC})=\rho _A\otimes \rho _B\). This is the product-marginal formulation in [ HJPW04 , Equations (5)–(7) ] .

Proof

The two relative entropies are

\begin{align} S(\rho _A)+S(\rho _{BC})-S(\rho _{ABC}), \notag \\ S(\rho _A)+S(\rho _B)-S(\rho _{AB}), \notag \end{align}

respectively. Cancelling the common term \(S(\rho _A)\) shows that their equality is precisely

\begin{align} S(\rho _{ABC})+S(\rho _B) & = S(\rho _{AB})+S(\rho _{BC}). \notag \end{align}

Entrywise, the partial-trace identity is

\begin{align} [\operatorname{tr}_C(\rho _A\otimes \rho _{BC})]_{(a,b),(a',b')} & = \sum _c(\rho _A)_{a,a'}(\rho _{BC})_{(b,c),(b',c)} \notag \\ & = (\rho _A)_{a,a'}(\rho _B)_{b,b'} = [\rho _A\otimes \rho _B]_{(a,b),(a',b')}. \notag \end{align}

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity. The raw Petz support map for the reference \(\rho _A\otimes \rho _{BC}\) recovers \(\rho _{ABC}\) from \(\rho _{AB}\). Moreover,

\begin{align} \rho _{ABC} & = (\operatorname {id}_A\otimes \mathcal R^{\mathrm{raw}}_{\rho _{BC}})(\rho _{AB}) = (\operatorname {id}_A\otimes \widehat{\mathcal R}_{\rho _{BC}})(\rho _{AB}), \notag \end{align}

where \(\widehat{\mathcal R}_{\rho _{BC}}\) is a trace-preserving completely positive extension of the support formula. This is equation (11) of [ HJPW04 ] , obtained from Theorem 3 and equation (8) together with the supported form of equation (10).

No marginal is required to be invertible. If \(\rho _A\) is singular, the displayed factorization is asserted on the supported input \(\rho _{AB}\) only; it is not a global factorization of an ambient product-reference completion.

Proof

Equality in strong subadditivity gives equality in relative-entropy data processing for \(\rho _{ABC}\) and \(\rho _A\otimes \rho _{BC}\). Lemma  13.6.47 supplies the required support inclusion, and Theorem 13.6.45 gives raw recovery after the canonical reassociation of the tensor factors.

The state \(\rho _{AB}\) is fixed on both sides by \(P_A\otimes \mathbf1_B\), so the supported product-reference factorization removes the first-factor support compression. It is also fixed by \(\mathbf1_A\otimes P_B\); hence every \(B\)-block lies in the support of \(\rho _B\). The completed local channel therefore agrees with its raw support formula on every block, proving both displayed identities.

For a finite-dimensional system \(A\), there is a finite family of effects \((M_s)_s\), with \(0\leq M_s\leq \mathbf1_A\), whose complex linear span is the full matrix algebra on \(A\). One member is the identity. For an operator \(X\) on \(A\otimes B\), define its conditional slice by

\begin{align} \xi _s(X) & = \operatorname{tr}_A\! \left(X(M_s\otimes \mathbf1_B)\right). \notag \end{align}

The family may be chosen from the four rank-one effects occurring in the polarization identity, scaled so that every member is bounded by the identity.

If two operators \(X,Y\) on \(A\otimes B\) have

\begin{align} \xi _s(X)=\xi _s(Y) \end{align}

for every member of the finite effect family, then \(X=Y\).

Proof

The polarization identity expresses every matrix unit as a complex linear combination of the selected rank-one effects. Hence any linear map vanishing on all selected effects vanishes on the full matrix algebra. Apply this to each matrix entry of the difference of the two conditional slice maps.

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and put \(\rho _{AB}=\operatorname{tr}_C\rho _{ABC}\). If \(\widehat{\mathcal R}:B\to B\otimes C\) is the completed recovery channel, define

\begin{align} \varphi & = \operatorname{tr}_C\circ \widehat{\mathcal R}. \end{align}

Then \(\varphi \) is trace-preserving and completely positive and

\begin{align} (\operatorname {id}_A\otimes \varphi )(\rho _{AB}) & = \rho _{AB}. \end{align}

For each selected effect, let

\begin{align} \xi _s & = \operatorname{tr}_A\! \left(\rho _{AB}(M_s\otimes \mathbf1_B)\right), & p_s & = \operatorname{tr}(\xi _s). \end{align}

The indices with \(p_s\neq 0\) form a finite nonempty set. For each such index, positivity of \(\xi _s\) makes \(p_s\) real and positive, and

\begin{align} \mu _s=p_s^{-1}\xi _s \end{align}

is a density operator and \(\varphi (\mu _s)=\mu _s\). Consequently a Kraus representation of \(\varphi \) gives a single preserving operation for this finite nonempty density family.

This is the conditional-family construction in [ HJPW04 , Theorem 6, lines 493–505 ] . Effects with \(p_s=0\) are omitted before normalization; no artificial conditional state is introduced for them.

Proof

Compose equation (11) with the partial trace over \(C\) to obtain the fixed point identity for \(\rho _{AB}\). Conditional slicing commutes with every linear map on \(B\), so \(\varphi (\xi _s)=\xi _s\). Positivity of \(\rho _{AB}\) and \(M_s\) gives positivity of \(\xi _s\). On the nonzero-trace subtype, scaling by \(p_s^{-1}\) preserves positivity and makes the trace one; the fixed-point equation is preserved by the same scaling. The identity effect has probability \(\operatorname{tr}(\rho _{AB})=1\), proving nonemptiness.

Let \(\rho _{ABC}\) be a tripartite density matrix and set \(\sigma _{ABC}=(\mathbf1_A/d_A)\otimes \rho _{BC}\). Then equality holds in strong subadditivity if and only if relative-entropy data processing under the partial trace over \(C\) is saturated for this pair:

\begin{align} D(\rho _{ABC} \| \sigma _{ABC}) & = D(\operatorname{tr}_C\rho _{ABC} \| \operatorname{tr}_C\sigma _{ABC}). \notag \end{align}

By Lemma 13.6.3, the reference on the right is \((\mathbf1_A/d_A)\otimes \rho _B\).

This is an equivalent hypothesis-free criterion with a different reference state. The exact product-marginal formulation is Theorem 13.6.48.

Proof

The two relative entropies are

\begin{align} \log d_A+S(\rho _{BC})-S(\rho _{ABC}), \notag \\ \log d_A+S(\rho _B)-S(\rho _{AB}), \notag \end{align}

respectively. The marginal support lemma places both singular references in the relative-entropy domain. Cancelling the common \(\log d_A\) shows that equality of the two relative entropies is precisely

\begin{align} S(\rho _{ABC})+S(\rho _B) & = S(\rho _{AB})+S(\rho _{BC}). \notag \end{align}
Theorem 13.6.54 SSA equality under factor relabeling

Let \(\rho \) be a Hermitian operator on \(A\otimes B\otimes C\), and let \(e_A:A'\to A\), \(e_B:B'\to B\), and \(e_C:C'\to C\) be bijections. Define \(\rho '\) by

\begin{align} \rho ’_{(a',b',c'),(\widetilde a',\widetilde b',\widetilde c')} & = \rho _{(e_A(a'),e_B(b'),e_C(c')), (e_A(\widetilde a'),e_B(\widetilde b'),e_C(\widetilde c'))}. \notag \end{align}

If \(\rho \) satisfies equality in strong subadditivity, then so does \(\rho '\).

Proof

The four operators entering the equality are related by the induced bijections:

\begin{align} \rho ’ & =\rho _{e_A\times e_B\times e_C}, \notag \\ \rho ’_{AB} & =(\rho _{AB})_{e_A\times e_B}, \notag \\ \rho ’_{BC} & =(\rho _{BC})_{e_B\times e_C}, \notag \\ \rho ’_B & =(\rho _B)_{e_B}. \notag \end{align}

These identities follow by changing variables in the three finite sums defining the partial traces. Entropy invariance under reindexing makes the corresponding four entropy terms equal, so the strong-subadditivity equality for \(\rho \) gives the equality for \(\rho '\).

Definition 13.6.55 Quantum Markov decomposition
#

A Hayashi Markov decomposition of a tripartite state \(\rho _{ABC}\) consists of a finite direct-sum decomposition

\begin{align} H_B & \cong \bigoplus _j H_{B_j^L}\otimes H_{B_j^R}, \notag \end{align}

together with a unitary change of basis on \(B\), a probability vector \((p_j)_j\), and density matrices \(\rho _{A B_j^L}\) and \(\rho _{B_j^R C}\) such that, in the adapted basis, the state becomes

\begin{align} \bigoplus _j p_j\, \rho _{A B_j^L}\otimes \rho _{B_j^R C}. \notag \end{align}

The terminology follows Hayashi’s presentation of quantum Markov structure  [ Hay06 ] ; the block decomposition used by the MPDO argument is the structure theorem of [ HJPW04 ] .

The finite-dimensional operator-algebraic step of the Hayden–Jozsa–Petz–Winter derivation of the Koashi–Imoto theorem ( [ HJPW04 , Appendix A ] ) begins with the common invariant algebra of a family of jointly invariant states.

Definition 13.6.56 Preserving operations and common average

Let \(\rho _1,\ldots ,\rho _K\) be density matrices in \(M_{D}(\mathbb {C})\), \(K\ge 1\). A trace-preserving completely positive Kraus family \(F\) preserves \(\rho _1,\ldots ,\rho _K\) if \(F\rho _k=\rho _k\) for every \(k\). Write

\begin{align} \mathbf F & =\{ F:\forall k,\ F\rho _k=\rho _k\} \label{eq:koashi_imoto_preserving_set} \end{align}

for the set of such operations – non-empty since the identity operation belongs to it – and

\begin{align} \bar\rho & =\frac1K\sum _{k=1}^K\rho _k \label{eq:koashi_imoto_common_average} \end{align}

for their common average.

Lemma 13.6.57 A preserving operation fixes the common average

Every \(F\in \mathbf F\) satisfies \(F\bar\rho =\bar\rho \).

Proof

Linearity turns \(F\rho _k=\rho _k\) for every \(k\) into \(F\bar\rho =\frac1K\sum _kF\rho _k=\frac1K\sum _k\rho _k=\bar\rho \).

Definition 13.6.58 Common invariant algebra

Assume in addition that the common average \(\bar\rho \) of Definition 13.6.56 is positive definite. Hayden–Jozsa–Petz–Winter instead reduce to this case by shrinking to the joint support of \(\rho _1,\ldots ,\rho _K\); that reduction is not re-derived here. By Lemma 13.6.57, \(\bar\rho \) is a positive definite fixed point of every \(F\in \mathbf F\), so Theorem 10.2.10 makes the fixed-point set of each adjoint map,

\begin{align} A_F & =\{ X:F^*(X)=X\} , \notag \end{align}

a \(*\)-subalgebra of \(M_{D}(\mathbb {C})\). The common invariant algebra is

\begin{align} A_0 & =\bigcap _{F\in \mathbf F}A_F, \label{eq:koashi_imoto_common_invariant_algebra} \end{align}

and \(X\in A_0\) if and only if \(F^*(X)=X\) for every \(F\in \mathbf F\).

Theorem 13.6.59 Block form of the common invariant algebra

Under the hypotheses of Definition 13.6.58, there are \(L\in \mathbb {N}\), positive dimensions \(d_0,\ldots ,d_{L-1}\) and multiplicities \(m_0,\ldots ,m_{L-1}\) with \(\sum _\ell d_\ell m_\ell =D\), and a unitary \(U\in M_{D}(\mathbb {C})\) such that a matrix \(A\in M_{D}(\mathbb {C})\) belongs to \(A_0\) exactly when

\begin{align} U^\dagger A U & =\bigoplus _{\ell =0}^{L-1}\mathbb {1}_{m_\ell }\otimes B_\ell \notag \end{align}

for some matrices \(B_\ell \in M_{d_\ell }(\mathbb {C})\).

Proof

Direct specialization of Theorem 10.6.9 to \(S=A_0\).

Theorem 13.6.60 Finite realization of the common invariant algebra

Under the hypotheses of Definition 13.6.58, there are finitely many operations \(F_1,\ldots ,F_M\in \mathbf F\), \(M\ge 1\), with

\begin{align} A_0=A_{F_1}\cap \cdots \cap A_{F_M}. \notag \end{align}
Proof

Because \(M_{D}(\mathbb {C})\) is finite-dimensional, among the finite intersections \(\bigcap _{F\in S}A_F\) over finite sets \(S\subseteq \mathbf F\) containing a fixed operation, one of minimal dimension already equals \(A_0\): enlarging \(S\) by one more operation either leaves the intersection unchanged or strictly decreases its dimension, and dimension cannot decrease indefinitely.

Theorem 13.6.61 A single operation realizes the common invariant algebra

Under the hypotheses of Definition 13.6.58, there is \(F_0\in \mathbf F\) with \(A_0=A_{F_0}\).

Proof

Take \(F_1,\ldots ,F_M\in \mathbf F\) as in Theorem 13.6.60 and pool their Kraus operators, each rescaled by \(1/\sqrt M\), into a single Kraus family representing

\begin{align} F_0=\frac1M\sum _{\mu =1}^M F_\mu . \notag \end{align}

Rescaling every pooled Kraus operator by \(1/\sqrt M\) turns the sum of the \(M\) trace-preservation identities into a single one, so \(F_0\in \mathbf F\); rescaling a Kraus operator by a nonzero scalar does not change the operators it commutes with, so \(A_{F_0}=A_{F_1}\cap \cdots \cap A_{F_M}=A_0\).

Definition 13.6.62 The projection onto the common invariant algebra

Fix \(F_0\in \mathbf F\) with \(A_0=A_{F_0}\) as in Theorem 13.6.61. Let \(P_0\) be the finite-dimensional mean-ergodic projection of the Schrödinger map \(F_0\), and define \(P_0^*\) as its trace-pairing adjoint.

For every \(X\in M_{D}(\mathbb {C})\),

\begin{align} \lim _{N\to \infty }\frac1N \sum _{n=0}^{N-1}(F_0^*)^n(X) =P_0^*(X). \label{eq:koashi_imoto_heisenberg_cesaro_convergence} \end{align}
Proof

The trace pairing transports every iterate of \(F_0^*\) to the corresponding iterate of \(F_0\), and hence transports each finite Cesàro average. Apply Theorem 10.1.5 to \(F_0\) and use nondegeneracy of the trace pairing.

Proof

Theorem 10.1.9 applied to \(F_0\) gives that \(P_0^*\) is positive, unital, and idempotent, with range equal to the fixed-point set of \(F_0^*\), which is \(A_{F_0}=A_0\).

Theorem 13.6.65 Simultaneous algebraic block form of the invariant family

Under the hypotheses of Definition 13.6.58, there are \(L\in \mathbb {N}\), positive dimensions \(d_\ell ,m_\ell \), a unitary \(U\), and density matrices \(\sigma _\ell \in M_{m_\ell }(\mathbb {C})\) such that, for every member \(\rho _k\) of the invariant family, there are matrices \(X_{\ell ,k}\in M_{d_\ell }(\mathbb {C})\) satisfying

\begin{align} U^\dagger \rho _kU & = \bigoplus _{\ell =0}^{L-1}\sigma _\ell \otimes X_{\ell ,k}. \label{eq:koashi_imoto_family_fixed_point_block_form} \end{align}

The density matrices \(\sigma _\ell \) and the decomposition are common to the whole family.

No positivity or trace normalization is asserted for the matrices \(X_{\ell ,k}\). Thus this is the full-support fixed-point block form underlying HJPW Appendix A, lines 853–856, not the normalized decomposition of Property 1.

Proof

Apply Theorem 10.8.16 to the Schrödinger map of the single operation \(F_0\) from Theorem 13.6.61. Its adjoint satisfies the Schwarz inequality by the Kraus form and trace preservation, and it fixes the positive-definite common average. Since \(F_0\rho _k=\rho _k\) for every \(k\), each family member belongs to the one fixed-point space \(U(\bigoplus _\ell \sigma _\ell \otimes M_{d_\ell }(\mathbb {C}))U^\dagger \).

Theorem 13.6.66 Normalized state decomposition of a full-support invariant family

Suppose in addition that every \(\rho _k\) is positive semidefinite with trace one. Under the explicit positive-definite common-average hypothesis of Definition 13.6.58, the common decomposition of Theorem 13.6.65 admits numbers \(q_{\ell |k}\geq 0\) and density matrices \(\tau _{\ell |k}\in M_{d_\ell }(\mathbb {C})\) such that, for every \(k\),

\begin{align} \sum _{\ell =0}^{L-1}q_{\ell |k}& =1, \label{eq:koashi_imoto_family_weights}\\ U^\dagger \rho _kU & = \bigoplus _{\ell =0}^{L-1} q_{\ell |k}\, \sigma _\ell \otimes \tau _{\ell |k}. \label{eq:koashi_imoto_normalized_family_block_form} \end{align}

The density matrix \(\sigma _\ell \) is independent of \(k\). The tensor factors are written in the reverse order from HJPW Property 1. This is the full-support specialization of that state decomposition. No action of a preserving operation on the summands is asserted in this theorem; the next theorem supplies it under the same full-support hypothesis.

Proof

Write the \(\ell \)-th block as \(\sigma _\ell \otimes X_{\ell ,k}\). Positivity of \(\rho _k\), compression to this block, and the partial trace over the first factor show that \(X_{\ell ,k}\) is positive semidefinite. Put \(q_{\ell |k}=\operatorname{tr}(X_{\ell ,k})\). The trace-one identities for \(\rho _k\) and \(\sigma _\ell \) give \(q_{\ell |k}\geq 0\) and \(\sum _\ell q_{\ell |k}=1\). For positive weight, divide \(X_{\ell ,k}\) by its trace. At zero weight, positivity forces \(X_{\ell ,k}=0\), so the maximally mixed density matrix may be chosen. In both cases \(X_{\ell ,k}=q_{\ell |k}\tau _{\ell |k}\).

Theorem 13.6.67 Action of preserving operations on full-support invariant-family blocks

Use the decomposition of Theorem 13.6.66. For every trace-preserving completely positive operation \(F\) satisfying \(F(\rho _k)=\rho _k\) for all \(k\), and every summand \(\ell \), there is a trace-preserving completely positive operation \(F_\ell \) on \(M_{m_\ell }(\mathbb {C})\) such that \(F_\ell (\sigma _\ell )=\sigma _\ell \). If \(\iota _\ell \) denotes inclusion of the \(\ell \)-th diagonal summand and

\begin{align} \widetilde F(X)=U^\dagger F(UXU^\dagger )U, \end{align}

then, for all \(A\in M_{m_\ell }(\mathbb {C})\) and \(B\in M_{d_\ell }(\mathbb {C})\),

\begin{align} \widetilde F\bigl(\iota _\ell (A\otimes B)\bigr) = \iota _\ell \bigl(F_\ell (A)\otimes B\bigr). \end{align}

This is the full-support specialization of HJPW Property \(2'\) (Appendix A, lines 808–816 and 860–882), with the tensor factors reversed. It does not include the joint-support reduction or the Stinespring form of Property 2.

Proof

For a preserving operation \(F\), every Kraus operator commutes with the common invariant algebra. In the common direct-sum coordinates it therefore has the form \(\bigoplus _\ell C_{i,\ell }\otimes \mathbf1\). Compressing \(\sum _iC_{i,\ell }^\dagger C_{i,\ell }=\mathbf1\) to a summand gives trace preservation of \(F_\ell \). The displayed action follows by multiplication. Positive definiteness of the common average implies that every summand has positive weight for some \(\rho _k\). Compressing \(F(\rho _k)=\rho _k\) and tracing over the second factor then gives \(F_\ell (\sigma _\ell )=\sigma _\ell \).

Theorem 13.6.68 Joint-support compression of an invariant state family

Let \(\rho _1,\ldots ,\rho _K\) be density operators and let \(Q\) be the support projection of their average. There are an integer \(r\) and an isometry \(V:\mathbb C^r\to \mathcal H\) with \(VV^\dagger =Q\) such that the compressed states \(\widehat\rho _k=V^\dagger \rho _kV\) are density operators, their average is positive definite, and

\begin{align} \rho _k=V\widehat\rho _kV^\dagger . \end{align}

Every trace-preserving completely positive operation preserving all \(\rho _k\) compresses along the same isometry to a trace-preserving completely positive operation \(\widehat F\) preserving all \(\widehat\rho _k\). Its ambient and compressed actions intertwine:

\begin{align} F(VXV^\dagger )=V\widehat F(X)V^\dagger . \end{align}

This is the reduction to the minimum joint supporting subspace in HJPW, Appendix A, lines 761–763.

Proof

The kernel of the average is the intersection of the kernels of the positive summands. Hence \(Q\rho _kQ=\rho _k\) for every \(k\). Choose an isometry onto the range of \(Q\). Compression then preserves positivity and trace, reconstructs each state, and makes the compressed average positive definite. Invariance of the average implies that every Kraus operator preserves its support. Compressing the Kraus operators therefore preserves both trace and every compressed state. Expanding the two Kraus sums and using support invariance gives the displayed intertwining identity.

Theorem 13.6.69 Invariant-family blocks on the joint support

Let \(\rho _1,\ldots ,\rho _K\) be density operators, without a faithfulness assumption on their average. On their minimum joint supporting subspace there is one direct-sum tensor decomposition in which

\begin{align} \widehat\rho _k = \bigoplus _j q_{j|k}\sigma _j\otimes \tau _{j|k}, \end{align}

where the weights form probability distributions and all displayed factors are density operators. Every trace-preserving completely positive operation preserving the compressed family acts on each diagonal summand as

\begin{align} F_j\otimes \operatorname {id}, \qquad F_j(\sigma _j)=\sigma _j. \end{align}

In particular, every operation preserving the original family restricts to such an operation, and its action transports back through the support isometry by the intertwining identity above. The support isometry reconstructs every original \(\rho _k\) from \(\widehat\rho _k\).

This is HJPW Properties 1 and \(2'\) after the joint-support reduction in Appendix A, lines 761–816 and 853–882. The tensor factors are written in the reverse order from HJPW.

Proof

Apply Theorem 13.6.68. The common average is positive definite in the resulting support coordinates, so Theorem 13.6.67 applies there. Apply the full-support block-action theorem to every preserving operation of the compressed family. Operations preserving the original family restrict along the same support isometry, and the intertwining identity transports their action back to the ambient space. Use the reconstruction identity for the family members.

Theorem 13.6.70 Invariant conditional blocks on the joint support

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and let \(\mu _s\) be the finite nonempty family of normalized conditional states obtained from the active separating effects. On the minimum joint supporting subspace of this family there are tensor-product direct-sum coordinates such that

\begin{align} \widehat\mu _s = \bigoplus _j q_{j|s}\sigma _j\otimes \tau _{j|s}, \end{align}

where the weights form probability distributions and all factors are density operators. The support isometry reconstructs every \(\mu _s\).

The channel

\begin{align} \varphi = \operatorname{tr}_C\circ \widehat{\mathcal R} \end{align}

restricts to the same support and intertwines with its ambient action. In these coordinates its action on each diagonal summand is

\begin{align} \varphi _j\otimes \operatorname {id}, \qquad \varphi _j(\sigma _j)=\sigma _j. \end{align}

The tensor factors are written in the reverse order from HJPW.

This is the application of HJPW Theorem 6, lines 493–505, to Properties 1 and \(2'\) from Appendix A, lines 761–816 and 853–882.

Proof

Theorem 13.6.52 supplies a finite nonempty density family together with a trace-preserving completely positive operation fixing every member. Apply Theorem 13.6.69 to this family. Retain the support reconstruction, the normalized block equations, and the restriction and block-action clauses for the preserving operation.

Theorem 13.6.71 Markov bipartite blocks on the joint support

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and let \(V\) be the support isometry and \((e,U,\sigma _j)\) the direct-sum coordinates of Theorem 13.6.70. Put \(W=\mathbf1_A\otimes V\). Then

\begin{align} W\bigl(W^\dagger \rho _{AB}W\bigr)W^\dagger & ={}\rho _{AB}. \end{align}

In the same support coordinates there are positive, not necessarily normalized operators \(\omega _j\) such that

\begin{align} (\mathbf1_A\otimes U)^\dagger W^\dagger \rho _{AB}W (\mathbf1_A\otimes U) & ={}\bigoplus _j\sigma _j\otimes \omega _j, \\ \sum _j\operatorname{tr}\omega _j& =1. \end{align}

The tensor factors are written in the reverse order from HJPW. The block equation is on the minimum joint supporting subspace; it does not assert an ambient direct-sum equivalence on subsystem \(B\).

Scope restriction (HJPW Theorem 6, equation (14)): the displayed direct sum is restricted to the minimum joint supporting subspace, whereas equation (14), lines 499–502, decomposes the ambient subsystem \(B\). This restriction is documented in the TNLean paper-gap note [ con26q ] .

Proof

Rescale the normalized equations for the active separating effects to their unnormalized conditional slices. A positive slice associated with an inactive effect has trace zero and therefore vanishes. The finite separating family then reconstructs \(\rho _{AB}\) through \(\mathbf1_A\otimes V\). In the transformed support coordinates, take \(\omega _j\) to be the partial trace over the common factor of the \(j\)-th principal block. Positivity follows from positivity of principal blocks and of partial trace. Applying separation once more gives the displayed direct-sum equation, and taking traces gives the normalization.

Theorem 13.6.72 Ambient Markov bipartite blocks

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity. Then the ambient middle subsystem has a direct-sum tensor decomposition

\begin{align} \mathcal H_B = \bigoplus _j \mathcal H_{b_j^R}\otimes \mathcal H_{b_j^L}. \end{align}

In unitary coordinates adapted to this decomposition there are density matrices \(\sigma _j\) and positive, not necessarily normalized operators \(\omega _j\) such that

\begin{align} \rho _{AB} & ={}\bigoplus _j \sigma _j\otimes \omega _j, \\ \sum _j\operatorname{tr}\omega _j & =1. \end{align}

The tensor factors appear here in the order \(b_j^R\otimes b_j^L\), the reverse of HJPW’s \(b_j^L\otimes b_j^R\) order. Directions complementary to the minimum joint support form one-dimensional tensor sectors with zero \(\omega _j\).

This is [ HJPW04 , Theorem 6, equations (13)–(14) ] , lines 493–502.

Proof

Extend the support isometry, after composing it with the support block unitary, to a unitary on the ambient middle subsystem. Split the orthogonal complement into one-dimensional tensor sectors. Give each such sector its unique density matrix and the zero unnormalized conditional factor. The supported sectors retain the factors from Theorem 13.6.71. Lifting the unitary extension through subsystem \(A\) and combining the complementary zero block with the supported direct sum gives the displayed ambient equation.

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, with the ambient decomposition

\begin{align} \mathcal H_B = \bigoplus _j \mathcal H_{b_j^R}\otimes \mathcal H_{b_j^L} \end{align}

from Theorem 13.6.72. There are a finite-dimensional ancilla, a fixed pure ancilla vector, and one unitary \(U_{BCE}\) whose block-coordinate form is

\begin{align} U_{BCE} = \bigoplus _j U_j\otimes \mathbf1_{b_j^L}. \end{align}

On every supported sector, \(U_j\) dilates a trace-preserving completely positive map from \(b_j^R\) to \(b_j^R C\) that sends \(\sigma _j\) to a state whose \(b_j^R\) marginal is \(\sigma _j\). The resulting state is positive and has trace one. The displayed block form uses the order \(U_j\otimes \mathbf1_{b_j^L}\); in HJPW’s order it is \(\mathbf1_{b_j^L}\otimes U_j\).

The recovery operation determined by the chosen unitary agrees with the Petz recovery operation on every operator supported by the minimum joint supporting subspace, and

\begin{align} (\operatorname {id}_A\otimes \widehat{\mathcal R})(\rho _{AB}) = \rho _{ABC}. \end{align}

No equality of the two operations is asserted on the complementary ambient sectors.

This is [ HJPW04 , Theorem 6, equation (15) ] , lines 547–560, using Appendix A, Theorem 10, Property 2, lines 791–800, the equivalence with Property \(2'\) in lines 808–823, and the operation-level proof in lines 853–882.

Proof

First obtain the ambient decomposition from Theorem 13.6.72. Choose rectangular Kraus operators for the Petz recovery operation and slice them along subsystem \(C\). On the minimum joint supporting subspace, Property \(2'\) makes every slice block diagonal with the identity on \(b_j^L\), while the induced operation on \(b_j^R\) fixes \(\sigma _j\). Extend each sector Stinespring isometry from the same fixed pure ancilla vector to a unitary. Their direct sum gives the block-coordinate unitary; conjugating by the ambient change of basis gives the physical unitary. Equality of the two Stinespring isometries on the support proves agreement with the Petz operation there. Combining this agreement with the supported reconstruction of \(\rho _{AB}\) gives the displayed recovery identity.

Let \(\rho _{ABC}\) be a tripartite density matrix attaining equality in strong subadditivity, and let \(U_B\) and the factors \(\omega _j\) be those of Theorems 13.6.72 and 13.6.73. For each supported sector, put

\begin{align} \widehat\rho _j & =\widehat{\mathcal R}_j(\sigma _j). \notag \end{align}

This is a density matrix on \(\mathcal H_{b_j^R}\otimes \mathcal H_C\), where \(\widehat{\mathcal R}_j\) is the sector recovery operation. Read each middle-system summand in the HJPW order \(\mathcal H_{b_j^L}\otimes \mathcal H_{b_j^R}\), and set \(W=\mathbf1_A\otimes (U_B\otimes \mathbf1_C)\). Then

\begin{align} W^\dagger \rho _{ABC}W & =0_{\perp }\oplus \bigoplus _j \omega _j\otimes \widehat\rho _j. \label{eq:entropy_hjpw_ambient_tripartite_blocks} \end{align}

Thus every complementary block is zero, while the supported \(j\)th block, in the factor order \((A b_j^L)\otimes (b_j^R C)\), is \(\omega _j\otimes \widehat\rho _j\). The factors \(\omega _j\) remain unnormalized; no probabilities or normalized left factors are asserted.

This is the final substitution in [ HJPW04 , Theorem 6, equations (11), (14), and (15) ] , lines 562–570.

Proof

Equations (11) and (14) give

\begin{align} \rho _{ABC} & =(\operatorname {id}_A\otimes \widehat{\mathcal R})(\rho _{AB}), \notag \\ (\mathbf1_A\otimes U_B)^\dagger \rho _{AB} (\mathbf1_A\otimes U_B) & =0_{\perp }\oplus \bigoplus _j\sigma _j\otimes \omega _j. \notag \end{align}

The physical unitary in equation (15) is the conjugate of its block unitary by \(U_B\otimes \mathbf1_C\). Hence conjugation by \(W\) changes the first identity into the sectorwise action of the block recovery operation on the second identity. In HJPW factor order the supported input block is \(\omega _j\otimes \sigma _j\). On the complement the input block is zero. On a supported sector,

\begin{align} (\operatorname {id}_{A b_j^L}\otimes \widehat{\mathcal R}_j) (\omega _j\otimes \sigma _j) & =\omega _j\otimes \widehat{\mathcal R}_j(\sigma _j) =\omega _j\otimes \widehat\rho _j. \notag \end{align}

Taking the direct sum proves (177) with the stated orientation \(W^\dagger \rho _{ABC}W\).

Let \(\rho _{ABC}\) be a tripartite density matrix satisfying

\begin{align} S(\rho _{ABC})+S(\rho _B) & = S(\rho _{AB})+S(\rho _{BC}). \notag \end{align}

Then there are a decomposition

\begin{align} \mathcal H_B & \cong \bigoplus _j \mathcal H_{B_j^L}\otimes \mathcal H_{B_j^R}, \notag \end{align}

a unitary \(V_B\), probabilities \(p_j\), and density matrices \(\rho _{A B_j^L}\) and \(\rho _{B_j^R C}\) such that, for \(V=\mathbf1_A\otimes V_B\otimes \mathbf1_C\),

\begin{align} V\rho _{ABC}V^\dagger & = \bigoplus _j p_j\, \rho _{A B_j^L}\otimes \rho _{B_j^R C}, \notag \\ p_j& \geq 0, \notag \\ \sum _jp_j& =1. \notag \end{align}
Proof

By Theorem 13.6.74, put \(W=\mathbf1_A\otimes (U_B\otimes \mathbf1_C)\). Then

\begin{align} W^\dagger \rho _{ABC}W & =0_\perp \oplus \bigoplus _j\omega _j\otimes \widehat\rho _j. \notag \end{align}

where \(\omega _j\geq 0\), each supported sector output state \(\widehat\rho _j\) is positive with trace one, and \(\sum _j\operatorname{tr}\omega _j=1\). Reindex every middle-system fibre from the recovery order \(B_j^R\otimes B_j^L\) to the Hayashi order \(B_j^L\otimes B_j^R\), and choose the Hayashi middle unitary \(V_B=U_B^\dagger \). This gives the required orientation \(V\rho _{ABC}V^\dagger =W^\dagger \rho _{ABC}W\).

For every ambient sector set

\begin{align} p_j& =\operatorname{Re}\operatorname{tr}\omega _j. \notag \end{align}

Positivity of \(\omega _j\) gives \(p_j\geq 0\), and the total trace identity gives \(\sum _jp_j=1\). Total positive-semidefinite normalization gives a density matrix \(\overline\omega _j\) on \(A B_j^L\) and a density matrix \(\overline\rho _j\) on \(B_j^R C\) such that

\begin{align} \omega _j\otimes \widehat\rho _j & =p_j\, \overline\omega _j\otimes \overline\rho _j. \notag \end{align}

On every supported sector, \(\overline\rho _j=\widehat\rho _j\), and \(\overline\omega _j=\omega _j/p_j\) whenever \(p_j{\gt}0\). If a supported sector has \(p_j=0\), its left factor may be any density matrix because \(0\, \overline\omega _j\otimes \widehat\rho _j=0\). On a complementary sector, where both raw factors vanish, either normalized factor may be chosen. These choices are made only at this final boundary, and every weighted summand is unchanged. The displayed direct sum is therefore a quantum Markov decomposition in the sense of Definition 13.6.55.

A tripartite density matrix \(\rho _{ABC}\) that admits a quantum Markov decomposition on the middle subsystem \(B\) satisfies

\begin{align} S(\rho _{ABC})+S(\rho _B) & = S(\rho _{AB})+S(\rho _{BC}). \notag \end{align}
Proof

Write the state in the adapted basis as the block-diagonal direct sum \(\bigoplus _j p_j\, \rho _{A B_j^L}\otimes \rho _{B_j^R C}\). The von Neumann entropy of a weighted orthogonal direct sum is \(S(\bigoplus _j p_j\, \omega _j) =-\sum _j p_j\log p_j+\sum _j p_jS(\omega _j)\) by Theorems 13.4.84 and 13.4.83, and the entropy of a tensor product is additive, \(S(\omega \otimes \tau )=S(\omega )+S(\tau )\), by Theorem 13.4.82. Tracing out one tensor factor within each block, the three reduced states factor as block-diagonal direct sums,

\begin{align} \rho _{AB} & \cong \bigoplus _j p_j\, \rho _{AB_j^L}\otimes \rho _{B_j^R}, \notag \\ \rho _{BC} & \cong \bigoplus _j p_j\, \rho _{B_j^L}\otimes \rho _{B_j^R C}, \notag \\ \rho _B & \cong \bigoplus _j p_j\, \rho _{B_j^L}\otimes \rho _{B_j^R}, \notag \end{align}

with \(\rho _{ABC}\cong \bigoplus _j p_j\, \rho _{AB_j^L}\otimes \rho _{B_j^R C}\) itself. Writing \(H=-\sum _j p_j\log p_j\), the four entropies expand as

\begin{align} S(\rho _{ABC}) & = H+\sum _j p_j [S(\rho _{AB_j^L})+S(\rho _{B_j^R C})], \notag \\ S(\rho _{AB}) & = H+\sum _j p_j [S(\rho _{AB_j^L})+S(\rho _{B_j^R})], \notag \\ S(\rho _B) & = H+\sum _j p_j [S(\rho _{B_j^L})+S(\rho _{B_j^R})], \notag \\ S(\rho _{BC}) & = H+\sum _j p_j [S(\rho _{B_j^L})+S(\rho _{B_j^R C})]. \notag \end{align}

Both sides of the claimed identity equal

\begin{align} 2H+\sum _j p_j\bigl[ S(\rho _{AB_j^L}) +S(\rho _{B_j^R C}) +S(\rho _{B_j^L}) +S(\rho _{B_j^R}) \bigr], \notag \end{align}

so \(S(\rho _{ABC})+S(\rho _B)=S(\rho _{AB})+S(\rho _{BC})\). The basis change on \(B\) and the direct-sum reindexing leave every entropy term unchanged.

For a tripartite density matrix \(\rho _{ABC}\),

\begin{align} S(\rho _{ABC})+S(\rho _B) & = S(\rho _{AB})+S(\rho _{BC}) \notag \end{align}

holds if and only if \(\rho _{ABC}\) admits a quantum Markov decomposition on the middle subsystem \(B\).

Proof

The forward implication is Theorem 13.6.75; the reverse implication is Theorem 13.6.76.

13.7 Mutual information

Definition 13.7.1 Mutual information
#

The quantum mutual information of a bipartite state \(\rho _{AB}\) is

\begin{align} I(A{:}B) & = S(\rho _A) + S(\rho _B) - S(\rho _{AB}). \label{eq:entropy_mutual_information} \end{align}
Theorem 13.7.2 Mutual information is non-negative

For any bipartite density matrix \(\rho _{AB}\), \(I(A{:}B) \ge 0\).

Proof

Apply strong subadditivity (Theorem 13.6.4) with trivial \(B\) (one-dimensional middle system). The SSA inequality \(S(\rho _{ABC}) + S(\rho _B) \le S(\rho _{AB}) + S(\rho _{BC})\) reduces to subadditivity \(S(\rho _{AC}) \le S(\rho _A) + S(\rho _C)\), which gives \(I(A{:}C) \ge 0\).

13.8 Entropy formulations

This section states the entropy formulations used later in the development: von Neumann entropy, strong subadditivity, quantum Markov decomposition, and mutual information. These statements are cited from the standard entropy literature and supply the entropy-theoretic input for the later MPDO arguments.

Definition 13.8.1 Entropy formulation of von Neumann entropy
#

This formulation has the same value \(S(\rho ) = -\sum _i \lambda _i \log \lambda _i\) as in Definition 13.4.1.

Theorem 13.8.2 Entropy formulation of strong subadditivity

For any tripartite density matrix \(\rho _{ABC}\) on \(A \otimes B \otimes C\),

\begin{align} S(\rho _{ABC}) + S(\rho _B) & \le S(\rho _{AB}) + S(\rho _{BC}). \label{eq:entropy_strong_subadditivity} \end{align}

This formulation introduces no new axiom: it is the same strong-subadditivity statement as Theorem 13.6.4, which is proved there from Lieb concavity along the relative-entropy route [ LR73 ] .

Definition 13.8.3 Entropy formulation of quantum Markov decomposition
#

This is the entropy formulation of the Hayashi Markov decomposition from Definition 13.6.55.

Theorem 13.8.4 SSA equality iff quantum Markov decomposition

For any tripartite density matrix \(\rho _{ABC}\), equality in strong subadditivity holds if and only if \(\rho _{ABC}\) admits a quantum Markov decomposition on the middle subsystem \(B\).

This formulation introduces no new axiom: it is the same equality criterion as Theorem 13.6.77.

Definition 13.8.5 Entropy formulation of mutual information
#

This formulation has the same value \(I(A{:}B) = S(\rho _A) + S(\rho _B) - S(\rho _{AB})\) as in Definition 13.7.1.

For a tripartite density matrix \(\rho _{ABC}\) with \(\dim B = 1\), one has \(S(\rho _{ABC}) \le S(\rho _{AB}) + S(\rho _{BC})\).

Proof

Apply Theorem 13.8.2 with one-dimensional middle subsystem. The reduced middle state has trace \(1\) by Theorem 13.13.3, hence its entropy vanishes by Theorem 13.13.4.

13.9 Mutual information: monotonicity and area-law bound

The two inequalities below are the downstream MPDO-facing consequences of the entropy inequalities in this chapter. The monotonicity inequality is the strong-subadditivity content underlying the MPDO mutual-information monotonicity \(I_L \le I_{L+1}\) (arXiv:1606.00608, Proposition C.1); the elementary area-law bound is the single-site entropy bound underlying the MPDO area-law bound \(I_L \le 4\log D\) (arXiv:1606.00608, cited from the Wolf area-law bound). Both inequalities follow directly from strong subadditivity (Theorem 13.6.4) and the single-system \(S(\rho ) \le \log D\) bound; neither introduces a new axiom.

Theorem 13.9.1 Entropy nonnegativity for finite index sets

For any PSD Hermitian matrix \(\rho \) with \(\operatorname{tr}(\rho ) = 1\) on an arbitrary finite index set, \(S(\rho ) \ge 0\). This is the analog of Theorem 13.4.4 for arbitrary finite index sets: it does not require the index set to be \(\{ 0,\ldots ,D{-}1\} \), so it applies to bipartite density matrices on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B}\).

Proof

The eigenvalues of a PSD matrix are non-negative, and eigenvalues of a trace-\(1\) Hermitian matrix with non-negative eigenvalues are bounded above by \(1\) (each single eigenvalue is at most the total sum). Apply \(x\log x \le 0\) on \([0, 1]\) pointwise and sum.

For any tripartite density matrix \(\rho _{ABC}\) on \(A \otimes B \otimes C\),

\begin{align} I(A{:}B)_{\operatorname{tr}_C \rho _{ABC}} & \le I(A{:}BC)_{\rho _{ABC}}, \label{eq:entropy_mutual_information_monotone} \end{align}

where the left-hand side is the bipartite mutual information of the reduced state \(\rho _{AB} = \operatorname{tr}_C(\rho _{ABC})\) and the right-hand side is evaluated by expanding \(I(A{:}BC) = S(\rho _A) + S(\rho _{BC}) - S(\rho _{ABC})\).

Proof

Expanding both sides in entropy form and cancelling the common \(S(\rho _A)\) term, the inequality reduces to strong subadditivity \(S(\rho _{ABC}) + S(\rho _B) \le S(\rho _{AB}) + S(\rho _{BC})\). The bipartite \(B\)-reduced state of \(\rho _{AB}\) matches the tripartite \(B\)-reduced state \(\operatorname{tr}_{AC}(\rho _{ABC})\), ensuring the \(S(\rho _B)\) term is the one supplied by Theorem 13.8.2.

Theorem 13.9.3 Elementary area-law bound on mutual information

For a bipartite density matrix \(\rho _{AB}\) on \(\mathbb {C}^{d_A} \otimes \mathbb {C}^{d_B}\) with \(d_A, d_B \ge 1\) whose single-system reduced states are obtained by partial trace, \(I(A{:}B) \le \log d_A + \log d_B\).

Proof

The reduced states are density matrices because partial trace preserves positivity and trace. Combine \(S(\rho _A) \le \log d_A\) and \(S(\rho _B) \le \log d_B\) (Theorem 13.4.7) with \(S(\rho _{AB}) \ge 0\) (Theorem 13.9.1).

13.10 Data processing under local channels

Theorem 13.10.1 Data processing on the first tensor factor

Let \(\rho _{AB}\) be a bipartite density operator and let \(\Phi _A\) be a trace-preserving completely positive map whose input and output matrix algebras may have different dimensions. Then \(I(A':B)_{(\Phi _A\otimes \operatorname{id}_B)(\rho )}\leq I(A:B)_\rho \).

Proof

Choose a rectangular Stinespring isometry for \(\Phi _A\). Conjugating by \(W=V_A\otimes \operatorname{id}_B\) gives \(\omega _{A'EB}=W\rho _{AB}W^\dagger \). Since \(W=V_A\otimes \operatorname{id}_B\) and \(V_A^\dagger V_A=\operatorname{id}_A\), its marginals satisfy

\begin{align} \omega _{A'E} & =V_A\rho _A V_A^\dagger , \notag \\ \omega _B & =\rho _B. \notag \end{align}

Entropy is preserved by the isometry \(W\), and also by \(V_A\) on the first marginal. Therefore

\begin{align} S(\omega _{A'EB}) & =S(\rho _{AB}), \notag \\ S(\omega _{A'E}) & =S(\rho _A), \notag \\ S(\omega _B) & =S(\rho _B). \notag \end{align}

Hence \(I(A'E:B)_\omega =I(A:B)_\rho \). The defining Stinespring identity is \(\operatorname{tr}_E(\omega _{A'EB})=(\Phi _A\otimes \operatorname{id}_B)(\rho _{AB})\). Strong subadditivity in the form \(I(R:A')\leq I(R:A'E)\) proves the stated inequality with \(R=B\).

Theorem 13.10.2 Data processing on the second tensor factor

Let \(\rho _{AB}\) be a bipartite density operator and let \(\Psi _B\) be a trace-preserving completely positive map whose input and output matrix algebras may have different dimensions. Then \(I(A:B')_{(\operatorname{id}_A\otimes \Psi _B)(\rho )}\leq I(A:B)_\rho \).

Proof

Exchange the two tensor factors and apply Theorem 13.10.1.

Theorem 13.10.3 Data processing for bipartite mutual information

Let \(\rho _{AB}\) be a bipartite density operator and let \(\Phi _A\) and \(\Psi _B\) be trace-preserving completely positive maps whose input and output matrix algebras may have different dimensions. Then \(I(A':B')_{(\Phi _A\otimes \Psi _B)(\rho )}\leq I(A:B)_\rho \).

Proof

Apply the two one-sided inequalities successively.

13.11 Classical information and operator-Schmidt bounds

Definition 13.11.1 Joint probability distribution and marginals

Let \(X\) and \(Y\) be finite sets. A matrix \(P=(P_{x,y})_{x\in X,y\in Y}\) is a joint probability distribution if \(P_{x,y}\geq 0\) for every \(x\in X\) and \(y\in Y\), and

\begin{align} \sum _{x\in X}\sum _{y\in Y}P_{x,y} & =1. \notag \end{align}

Its row and column marginals are respectively

\begin{align} p_x & =\sum _{y\in Y}P_{x,y}, \notag \\ q_y & =\sum _{x\in X}P_{x,y}. \notag \end{align}
Definition 13.11.2 Entropy and classical mutual information

For \(t{\gt}0\), set \(h(t)=-t\log t\), and set \(h(0)=0\). The entropy of a probability distribution \(a=(a_z)_{z\in Z}\) on a finite set is \(H(a)=\sum _{z\in Z}h(a_z)\). For a joint probability distribution \(P\) with row and column marginals \(p\) and \(q\), put

\begin{align} H(P) & =\sum _{x\in X}\sum _{y\in Y}h(P_{x,y}), \notag \\ I(X:Y)_P & =H(p)+H(q)-H(P). \notag \end{align}
Lemma 13.11.3 Entropy is bounded by the logarithm of the support cardinality

Let \(p=(p_z)_{z\in Z}\) be a probability distribution on a finite set \(Z\). Define \(h(t)=-t\log t\) for \(t{\gt}0\) and \(h(0)=0\). Then

\begin{align} \sum _{z\in Z}h(p_z) & \leq \log \left\lvert \{ z\in Z:p_z\neq 0\} \right\rvert . \notag \end{align}
Proof

Let \(k=\lvert \{ z\in Z:p_z\neq 0\} \rvert \). Jensen’s inequality for the concave function \(h\), with uniform weights on the support, gives

\begin{align} \frac{1}{k}\sum _{\substack {z\in Z \\ p_z\neq 0}}h(p_z) & \leq h\left(\frac{1}{k} \sum _{\substack {z\in Z \\ p_z\neq 0}}p_z\right) =h\left(\frac{1}{k}\right) =\frac{\log k}{k}. \notag \end{align}

Multiplication by \(k\) proves the result, since values outside the support contribute \(h(0)=0\).

Theorem 13.11.4 Classical mutual information is bounded by ordinary rank

Let \(X\) and \(Y\) be finite sets and let \(P=(P_{x,y})_{x\in X,y\in Y}\) be a joint probability distribution. With natural logarithms, \(I(X:Y)_P\leq \log \operatorname{rank}_{\mathbb {R}}P\). Equivalently, with logarithms to base two, \(2^{I(X:Y)_P}\leq \operatorname{rank}_{\mathbb {R}}P\).

Proof

This is Theorem 4.1 of [ RV17 ] . If the number of nonzero rows exceeds the ordinary real rank, choose a nontrivial real linear relation among those rows. Nonnegativity forces the relation to have coefficients of both signs. Rescaling each row by \(1+\varepsilon \beta _x\) gives two endpoint distributions. For every \(y\in Y\), the row relation gives

\begin{align} \sum _x(1+\varepsilon \beta _x)P_{x,y} & =\sum _xP_{x,y} +\varepsilon \underbrace{\sum _x\beta _xP_{x,y}}_{=0} =\sum _xP_{x,y}. \notag \end{align}

Thus each endpoint preserves every column marginal, removes at least one row, and has no larger rank. Write \(h(t)=-t\log t\) for \(t{\gt}0\), with \(h(0)=0\), and put \(m=\min _x\beta _x{\lt}0{\lt}M=\max _x\beta _x\). Then

\begin{align} I(X:Y)_{P^{(\varepsilon )}} & =I(X:Y)_P+\varepsilon S(\beta ), \notag \\ S(\beta ) & =\sum _x\beta _x\left(h(p_x)-\sum _y h(P_{x,y})\right). \notag \end{align}

If \(S(\beta )\geq 0\), choose \(\varepsilon =-m^{-1}\geq 0\); otherwise choose \(\varepsilon =-M^{-1}\leq 0\). In either case \(\varepsilon S(\beta )\geq 0\), so the selected endpoint has mutual information no smaller than the original distribution. Repeating the construction leaves at most \(\operatorname{rank}_{\mathbb {R}}P\) nonzero rows. Finally, \(I(X:Y)\leq H(X)\), and the entropy of a distribution supported on \(k\) points is at most \(\log k\).

13.11.1 Operator-Schmidt rank and marginal-support compression

Theorem 13.11.1.1 Positivity of the operator-Schmidt rank
#

Every nonzero finite-dimensional bipartite operator \(X\) satisfies

\begin{align} 0 & {\lt}\operatorname {OSR}(X). \notag \end{align}

This includes zero-dimensional factors, where the nonzero hypothesis is impossible.

Proof

If \(\operatorname {OSR}(X)=0\), a shortest product decomposition of \(X\) is an empty sum. Hence \(X=0\).

Theorem 13.11.1.2 Coefficient-space characterization of operator-Schmidt rank

Write \(\rho _{ij}\in M_{d_B}(\mathbb {C})\) for the blocks determined by a basis of the first factor, and define \(\mathcal R_\rho (X)=\sum _{i,j}X_{ij}\rho _{ij}\). Then

\begin{align} \operatorname {OSR}(\rho ) & =\dim \operatorname{range}\mathcal R_\rho =\dim \operatorname{span}\{ \rho _{ij}:1\leq i,j\leq d_A\} . \notag \end{align}
Proof

The minimality theorem at the preceding definition proves that the range dimension \(\dim \operatorname{range}\mathcal R_\rho \) is the least admissible product-decomposition length: it is itself an admissible length and no shorter length suffices. Since the operator-Schmidt rank is defined as the least such length, uniqueness of the least integer identifies the two, giving the first equality \(\operatorname {OSR}(\rho )=\dim \operatorname{range}\mathcal R_\rho \). For the second equality, \(\mathcal R_\rho \) is the linear combination map sending a coefficient matrix \(X\) to \(\sum _{i,j}X_{ij}\rho _{ij}\), so its range is exactly the span of the operator blocks \(\rho _{ij}\).

Theorem 13.11.1.3 Operator-Schmidt rank under local isometries

Let \(V_A:\mathcal H_A\to \mathcal K_A\) and \(V_B:\mathcal H_B\to \mathcal K_B\) be isometries between finite-dimensional complex spaces. For every operator \(X\) on \(\mathcal H_A\otimes \mathcal H_B\),

\begin{align} \operatorname {OSR}\! \left( (V_A\otimes V_B)X(V_A\otimes V_B)^\dagger \right) & =\operatorname {OSR}(X). \notag \end{align}

This remains valid when one or more of the spaces have dimension zero.

Proof

More generally, multiplying on the left and right by product matrices cannot increase operator-Schmidt rank: apply the four local matrices to the two factors in each term of a shortest product decomposition. Applying this inequality first to \(V_A,V_B\) and then to their adjoints gives the two inequalities. The identities \(V_A^\dagger V_A=1\) and \(V_B^\dagger V_B=1\) identify the twice-transformed operator with \(X\).

Theorem 13.11.1.4 Support of a bipartite positive operator

Let \(\rho _{AB}\succeq 0\), with marginals \(\rho _A\) and \(\rho _B\). Then

\begin{align} \ker (\rho _A\otimes \rho _B) & \subseteq \ker \rho _{AB}. \notag \end{align}

Equivalently, the support of \(\rho _{AB}\) is contained in \(\operatorname{supp}(\rho _A)\otimes \operatorname{supp}(\rho _B)\). This remains valid when either factor has dimension zero.

Proof

Let \(P_A\) and \(P_B\) be the support projections of the marginals. The two right-absorption identities give \(\rho _{AB}(P_A\otimes P_B)=\rho _{AB}\). The support projection of \(\rho _A\otimes \rho _B\) is \(P_A\otimes P_B\). Therefore every vector annihilated by \(\rho _A\otimes \rho _B\) is annihilated by \(\rho _{AB}\).

Theorem 13.11.1.5 Compression to both marginal supports

Let \(\rho \geq 0\) be an operator on \(H_A\otimes H_B\). Let \(V_A:\widehat H_A\to H_A\) and \(V_B:\widehat H_B\to H_B\) be isometries whose range projectors are the support projectors \(P_A\) and \(P_B\) of the two marginals. Set \(W=V_A\otimes V_B\) and \(\rho _c=W^\dagger \rho W\). Then

\begin{align} W\rho _cW^\dagger & =\rho , & \operatorname {OSR}(\rho _c) & =\operatorname {OSR}(\rho ). \notag \end{align}

No marginal is required to be faithful, and zero-dimensional coordinate spaces are permitted.

Proof

The four marginal-support absorption identities give

\begin{align} (P_A\otimes P_B)\rho & =\rho , & \rho (P_A\otimes P_B) & =\rho . \notag \end{align}

Since \(WW^\dagger =P_A\otimes P_B\), associativity yields

\begin{align} W\rho _cW^\dagger & =(WW^\dagger )\rho (WW^\dagger )=\rho . \notag \end{align}

The local-isometry theorem applied to \(\rho _c\) gives \(\operatorname {OSR}(W\rho _cW^\dagger )=\operatorname {OSR}(\rho _c)\). Substitution of the reconstruction identity proves the rank equality.

Lemma 13.11.1.6 Partial traces under product isometries

Let \(A:H_A\to K_A\) and \(B:H_B\to K_B\) be linear maps between finite-dimensional complex spaces, and let \(X\) be an operator on \(H_A\otimes H_B\). If \(B^\dagger B=1\), then

\begin{align} \operatorname{tr}_{K_B}\! \left((A\otimes B)X(A\otimes B)^\dagger \right) & =A(\operatorname{tr}_{H_B}X)A^\dagger . \notag \end{align}

If \(A^\dagger A=1\), then, symmetrically,

\begin{align} \operatorname{tr}_{K_A}\! \left((A\otimes B)X(A\otimes B)^\dagger \right) & =B(\operatorname{tr}_{H_A}X)B^\dagger . \notag \end{align}

The identities include zero-dimensional spaces.

Proof

Expand the matrix entries in orthonormal bases. In the first identity, the sum over the traced output basis gives the matrix entries of \(B^\dagger B=1\), and hence collapses the two input indices on \(H_B\). The second identity follows by the same argument with the factors interchanged.

In the setting of Theorem 13.11.1.5, the two marginals of \(\rho _c\) satisfy

\begin{align} \operatorname{tr}_{\widehat H_B}\rho _c & =V_A^\dagger (\operatorname{tr}_{H_B}\rho )V_A, & \operatorname{tr}_{\widehat H_A}\rho _c & =V_B^\dagger (\operatorname{tr}_{H_A}\rho )V_B. \notag \end{align}

Both compressed marginals are positive definite, including when one of their spaces has dimension zero.

Proof

Insert the reconstruction \(\rho =W\rho _cW^\dagger \) into the two covariance identities of Lemma 13.11.1.6, and multiply by the corresponding adjoint isometries. Each displayed marginal is the compression of a positive semidefinite matrix to its support, so its positive definiteness follows from Theorem 10.4.1.5.

Let \(\rho \geq 0\) be a bipartite complex matrix whose first marginal is faithful. Choose an eigenbasis in which \(\operatorname{tr}_B\rho =\operatorname{diag}(p_1,\ldots ,p_{d_A})\) with every \(p_i{\gt}0\), and define the linear map \(\Phi _\rho \) on matrix units by

\begin{align} \Phi _\rho (E_{ij}) & =\frac{\rho _{ij}}{\sqrt{p_ip_j}}. \notag \end{align}

Write \(\sigma =\operatorname{diag}(p_1,\ldots ,p_{d_A})\) and \(\tau =\operatorname{tr}_A\rho \). Then \(\Phi _\rho \) is completely positive and trace preserving,

\begin{align} \Phi _\rho (\sigma ) & =\tau , & \dim \operatorname{range}\Phi _\rho & =\operatorname {OSR}(\rho ), \notag \end{align}

and application of \(\Phi _\rho \) to the first half of \(\sum _{i,j}\sqrt{p_ip_j}\, E_{ij}\otimes E_{ij}\), followed by restoring the order of the two factors, reconstructs \(\rho \).

Proof

Strict positivity of the \(p_i\) makes the entrywise input scaling invertible, so it does not change the range. Positivity of \(\rho \) gives a Kraus representation of \(\Phi _\rho \), while the first marginal equation gives trace preservation. Summing the diagonal blocks gives \(\Phi _\rho (\sigma )=\tau \). Substitution on matrix units proves the reconstruction identity. This is the finite-dimensional Choi representation [ Cho75 ] with the input marginal absorbed into the canonical purification.

13.12 Support compression for entropy functionals

Definition 13.12.1 Isometric conjugation into a larger matrix algebra

Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) be an isometry, so that \(V^\dagger V=\mathbb {1}_k\). Define \(\iota _V:M_{k}(\mathbb {C})\to M_{D}(\mathbb {C})\) by

\begin{align} \iota _V(A) & =VAV^\dagger . \label{eq:mpdo_isometric_conjugation} \end{align}

The map \(\iota _V\) is complex-linear, multiplicative, and \(*\)-preserving. In general it is not unital: \(\iota _V(\mathbb {1}_k)=VV^\dagger \) is the projection onto the range of \(V\).

Theorem 13.12.2 Functional calculus under rectangular isometric conjugation
#

Let \(A\in M_{k}(\mathbb {C})\) be Hermitian, let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy \(V^\dagger V=\mathbb {1}_k\), and let \(f:\mathbb {R}\to \mathbb {R}\) satisfy \(f(0)=0\). Then

\begin{align} f(VAV^\dagger ) & =Vf(A)V^\dagger , \label{eq:mpdo_cfc_isometric_conjugation} \end{align}

where both sides use the continuous functional calculus on the respective finite spectra.

Proof

Apply functoriality of the continuous functional calculus to \(\iota _V\). The condition \(f(0)=0\) removes the contribution from the orthogonal complement of the range of \(V\), where \(\iota _V(A)\) vanishes.

Theorem 13.12.3 Trace preservation under isometric inclusion

Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy \(V^\dagger V=\mathbb {1}_k\). For every \(A\in M_{k}(\mathbb {C})\),

\begin{align} \operatorname{tr}(VAV^\dagger ) & =\operatorname{tr}(A). \label{eq:mpdo_isometric_trace} \end{align}
Proof

Cyclicity of the trace gives \(\operatorname{tr}(VAV^\dagger )=\operatorname{tr}(V^\dagger VA)=\operatorname{tr}(A)\).

Theorem 13.12.4 Compression cancels isometric expansion

Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy \(V^\dagger V=\mathbb {1}_k\). For every \(A\in M_{k}(\mathbb {C})\),

\begin{align} V^\dagger (VAV^\dagger )V & =A. \label{eq:mpdo_isometric_compression_expansion} \end{align}
Proof

Associativity and \(V^\dagger V=\mathbb {1}_k\) reduce the left-hand side to \((V^\dagger V)A(V^\dagger V)=A\).

Theorem 13.12.5 Reconstruction from a reference support

Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite and suppose that \(\ker \omega \subseteq \ker \rho \). If \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfies \(VV^\dagger =P_{\operatorname{supp}(\omega )}\), then

\begin{align} V(V^\dagger \rho V)V^\dagger & =\rho . \label{eq:mpdo_support_reconstruction} \end{align}

No condition on \(V^\dagger V\) is needed for this identity.

Proof

The kernel inclusion and positivity give \(\rho P_{\operatorname{supp}(\omega )}=\rho \); taking adjoints also gives \(P_{\operatorname{supp}(\omega )}\rho =\rho \). Substituting \(VV^\dagger =P_{\operatorname{supp}(\omega )}\) on both sides of \(\rho \) proves (187).

Theorem 13.12.6 Relative entropy under support compression

Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite with \(\ker \omega \subseteq \ker \rho \). Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy

\begin{align} V^\dagger V & =\mathbb {1}_k, \\ VV^\dagger & =P_{\operatorname{supp}(\omega )}. \label{eq:mpdo_reference_support_isometry} \end{align}

Then

\begin{align} D(\rho \Vert \omega ) & =D(V^\dagger \rho V\Vert V^\dagger \omega V). \label{eq:mpdo_relative_entropy_support_compression} \end{align}

The logarithm is totalized by \(\log 0=0\) on the kernel.

Proof

Theorem 13.12.5 reconstructs both \(\rho \) and \(\omega \) from their compressions. Apply Theorem 13.12.2 to the logarithm, use multiplicativity of \(\iota _V\), and finish with (185).

If \(A\in M_{n}(\mathbb {C})\) is Hermitian and \(f:\mathbb {R}\to \mathbb {R}\), then

\begin{align} f(A)^T & =f(A^T). \label{eq:mpdo_cfc_transpose} \end{align}

In particular,

\begin{align} \log (A^T) & =(\log A)^T, & P_{A^T} & =P_A^T. \label{eq:mpdo_log_support_transpose} \end{align}

If \(A\) is positive semidefinite, then for every \(r\in \mathbb {R}\),

\begin{align} (A^T)^r & =(A^r)^T, & \sqrt{A^T} & =(\sqrt A)^T. \label{eq:mpdo_power_sqrt_transpose} \end{align}

These identities include singular matrices and matrices whose index set is empty. They are project-derived rather than statements of CPSV16.

Proof

For a Hermitian matrix, transpose is entrywise complex conjugation. Apply covariance of the continuous functional calculus under this real star-algebra automorphism. Continuity of \(f\) is needed only on the finite spectrum of \(A\), where it is automatic. The logarithm, real-power, and square-root identities follow by specialization. For the support projection, specialize to \(f(x)=1\) for \(x\neq 0\) and \(f(0)=0\).

Theorem 13.12.8 Logarithm of a positive square root
#

If \(\rho \succeq 0\), then

\begin{align} \log \sqrt{\rho } & =\frac12\log \rho . \label{eq:mpdo_log_sqrt} \end{align}
Proof

On every positive eigenvalue this is the scalar identity \(\log \sqrt{x}=\frac12\log x\). At \(x=0\), both sides vanish under the convention \(\log 0=0\).

Theorem 13.12.9 Logarithm of a real power
#

If \(A\succ 0\), then, for every \(s\in \mathbb {R}\),

\begin{align} \log (A^s) & =s\log A. \label{eq:mpdo_log_rpow} \end{align}
Proof

Every eigenvalue of \(A\) is positive, so the assertion follows from \(\log (x^s)=s\log x\) on \((0,\infty )\) and the continuous functional calculus.

Theorem 13.12.10 Support-aware vector-state Jensen inequality for the logarithm

Let \(A\in M_{n}(\mathbb {C})\) be positive semidefinite and let \(v\in \mathbb {C}^n\) satisfy \(\langle v,v\rangle =1\) and \(P_{\operatorname{supp}(A)}v=v\). Then

\begin{align} \operatorname{Re}\langle v,(\log A)v\rangle & \leq \log \! \left(\operatorname{Re}\langle v,Av\rangle \right). \label{eq:mpdo_support_log_jensen} \end{align}

Here \(\log A\) is defined by the continuous functional calculus with \(\log 0=0\). Compare the finite-dimensional vector-state Jensen argument in [ Bha97 , Chapter V ] .

Proof

Diagonalize \(A=U\operatorname{diag}(\mu _i)U^\dagger \), set \(w=U^\dagger v\), and put \(p_i=|w_i|^2\). The support condition implies \(\mu _i=0\Longrightarrow w_i=0\), so every spectral weight at zero vanishes. Thus \(p_i\geq 0\) for every \(i\in S=\{ i:\mu _i{\gt}0\} \) and \(\sum _{i\in S}p_i=1\). The spectral formulas give

\begin{align} \operatorname{Re}\langle v,(\log A)v\rangle & =\sum _{i\in S}p_i\log \mu _i, \notag \\ \operatorname{Re}\langle v,Av\rangle & =\sum _{i\in S}p_i\mu _i. \notag \end{align}

Scalar concavity of the logarithm on \((0,\infty )\) now gives (196). No concavity assertion at zero and no positive-definiteness of \(A\) are used.

Definition 13.12.11 Order-two sandwiched trace functional
#

For square matrices \(\rho \) and \(\omega \) of the same size, define

\begin{align} Q_2(\rho ,\omega ) & =\operatorname{Re}\operatorname{tr}\left[ \left(\omega ^{-1/4}\rho \, \omega ^{-1/4}\right)^2 \right]. \label{eq:mpdo_sandwiched_two_trace} \end{align}

The negative power is defined by functional calculus and vanishes on the kernel of \(\omega \). On trace-one positive semidefinite matrices satisfying \(\ker \omega \subseteq \ker \rho \), its logarithm is the order-two sandwiched Rényi divergence [ MLDS\(^{+}\)13 , Definition 2 ] .

Theorem 13.12.12 Nonnegativity of the order-two functional

If \(\rho \succeq 0\) and \(\omega \succeq 0\), then \(Q_2(\rho ,\omega )\geq 0\).

Proof

The sandwiched matrix \(\omega ^{-1/4}\rho \, \omega ^{-1/4}\) is positive semidefinite. The trace of its square is therefore nonnegative.

Theorem 13.12.13 Faithful cyclic form of the order-two functional

If \(\omega \succ 0\), then, for every square matrix \(\rho \) of the same size,

\begin{align} Q_2(\rho ,\omega ) & =\operatorname{Re}\operatorname{tr}\left(\rho \, \omega ^{-1/2} \rho \, \omega ^{-1/2}\right). \label{eq:mpdo_sandwiched_two_weighted} \end{align}
Proof

Combine the two adjacent inverse quarter-powers by \(\omega ^{-1/4}\omega ^{-1/4}=\omega ^{-1/2}\) and rotate the factors under the trace. No Hermiticity assumption on \(\rho \) is needed.

Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) satisfy \(\rho \succeq 0\), \(\omega \succ 0\), and \(\operatorname{tr}\rho =1\). Then

\begin{align} D(\rho \Vert \omega ) & \leq \log Q_2(\rho ,\omega ). \label{eq:mpdo_relative_q2_posdef} \end{align}

No normalization of \(\omega \) is required.

Proof

Set \(R=\sqrt\rho \), \(T=\omega ^{-1/2}\), \(\Delta =R^T\otimes T\), and \(v=\operatorname {vec}(R)\). The vectorization is column-stacking, with factor order

\begin{align} (B\otimes A)\operatorname {vec}(X) & =\operatorname {vec}(AXB^T). \label{eq:mpdo_column_vec_kronecker} \end{align}

Hence \(\Delta v=\operatorname {vec}(T\rho )\), and cyclicity of the trace identifies its first moment with

\begin{align} m=\operatorname{Re}\langle v,\Delta v\rangle & =\langle R,RTR\rangle _{\mathrm{HS}}. \label{eq:mpdo_relative_modular_half_moment} \end{align}

The support projection of \(\Delta \) fixes \(v\):

\begin{align} P_{\operatorname{supp}(\Delta )}v & =(P_{\operatorname{supp}(R^T)}\otimes P_{\operatorname{supp}(T)})\operatorname {vec}(R) =\operatorname {vec}(R)=v. \label{eq:mpdo_relative_modular_support_fixing} \end{align}

Indeed, \(T\) is positive definite, so \(P_{\operatorname{supp}(T)}=1\), while the support projection of \(R^T\) fixes \(R^T\) and hence \(R P_{\operatorname{supp}(R^T)}^T=R\). The column-stacking identity (200) proves the displayed equality. This support condition is essential when \(\rho \) is singular: it removes the zero spectral weights before Jensen’s inequality is applied.

Theorem 13.4.70, 13.12.8, 13.12.9, and 13.12.7 give

\begin{align} \operatorname{Re}\langle v,(\log \Delta )v\rangle & =\frac12 D(\rho \Vert \omega ). \label{eq:mpdo_relative_modular_log_moment} \end{align}

Theorem 13.12.10 therefore yields \(\frac12D(\rho \Vert \omega )\leq \log m\). Since \(\| R\| _{\mathrm{HS}}^2=\operatorname{tr}\rho =1\), Hilbert–Schmidt Cauchy–Schwarz gives

\begin{align} m^2 & \leq \| R\| _{\mathrm{HS}}^2\| RTR\| _{\mathrm{HS}}^2 =Q_2(\rho ,\omega ). \label{eq:mpdo_relative_modular_cauchy_schwarz} \end{align}

Positivity of \(m\) now gives \(D(\rho \Vert \omega )\leq \log (m^2)\leq \log Q_2(\rho ,\omega )\).

The auxiliary divergence and its vector-state Jensen argument are adapted from [ MLDS\(^{+}\)13 , arXiv:1306.3142v4, Definition 5 and Lemma 19, used in the proof of Theorem 7 ] . Only the direct order-two endpoint is used here; the final Cauchy–Schwarz estimate is not an application of Beigi’s interpolation theorem.

Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite with \(\ker \omega \subseteq \ker \rho \). Let \(V:\mathbb {C}^k\to \mathbb {C}^D\) satisfy (189). Then

\begin{align} Q_2(\rho ,\omega ) & =Q_2(V^\dagger \rho V,V^\dagger \omega V). \label{eq:mpdo_sandwiched_two_support_compression} \end{align}
Proof

The function \(x\mapsto x^{-1/4}\) is taken to be zero at \(x=0\), so Theorem 13.12.2 applies. Reconstruct \(\rho \) and \(\omega \), combine the three expanded factors by multiplicativity of \(\iota _V\), and use trace preservation.

Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite matrices of trace one with \(\ker \omega \subseteq \ker \rho \). Assume that for every dimension and every pair of trace-one matrices \(A\succeq 0\) and \(B\succ 0\),

\begin{align} D(A\Vert B) & \leq \log Q_2(A,B). \label{eq:mpdo_faithful_relative_q2_hypothesis} \end{align}

Then

\begin{align} D(\rho \Vert \omega ) & \leq \log Q_2(\rho ,\omega ). \label{eq:mpdo_singular_relative_q2_comparison} \end{align}

Thus this theorem reduces the singular-reference case to the faithful comparison. Theorem 13.12.14 supplies the hypothesis in (206).

Proof

Choose an isometric inclusion of the support of \(\omega \). Its compression \(\omega _V=V^\dagger \omega V\) is positive definite by Theorem 10.4.1.5, while \(\rho _V=V^\dagger \rho V\) remains positive semidefinite. Reconstruction and (185) show that both compressed matrices have trace one. Apply (206) to \((\rho _V,\omega _V)\), then use (190) and (205).

Theorem 13.12.17 Relative entropy bounded by the support-restricted order-two functional

Let \(\rho ,\omega \in M_{D}(\mathbb {C})\) be positive semidefinite matrices of trace one. If \(\ker \omega \subseteq \ker \rho \), then

\begin{align} D(\rho \Vert \omega ) & \leq \log Q_2(\rho ,\omega ). \label{eq:mpdo_relative_q2_support} \end{align}
Proof

Apply Theorem 13.12.16 with the faithful comparison from Theorem 13.12.14. Compression to the support of \(\omega \) preserves both entropy functionals and both trace normalizations, and makes the reference positive definite.

13.12.1 Whitened Choi estimates and sandwiched Rényi bounds

Let \(\rho _{AB}\geq 0\), and choose an eigenbasis of its faithful first marginal \(\sigma =\operatorname{diag}(p_1,\ldots ,p_{d_A}){\gt}0\). Suppose also that the second marginal \(\tau \) is positive definite. For the supported-marginal channel \(\Phi _\rho \), set

\begin{align} L(X) & =\tau ^{-1/4} \Phi _\rho (\sigma ^{1/4}X\sigma ^{1/4})\tau ^{-1/4}, \\ W & =(\sigma ^T\otimes \tau )^{-1/4}\rho _{AB} (\sigma ^T\otimes \tau )^{-1/4}. \notag \end{align}

Then

\begin{align} W & =J(L), & \dim _{\mathbb {C}}\operatorname{range}L & =\dim _{\mathbb {C}}\operatorname{range}\Phi _\rho . \notag \end{align}

Here \(\sigma ^T=\sigma \) because the chosen basis diagonalizes \(\sigma \).

Proof

For arbitrary congruence matrices \(A\) and \(B\), the matrix-unit convention for the Choi matrix gives

\begin{align} J(\Gamma _B\circ \Phi \circ \Gamma _A) & =\Gamma _{A^T\otimes B}(J(\Phi )). \notag \end{align}

Apply this identity to the inverse-square-root scaling in \(\Phi _\rho \) and to the two fourth-root factors defining \(L\). Entrywise diagonal functional calculus gives \(\sigma ^{1/4}=\operatorname{diag}(p_i^{1/4})\) and \(\operatorname{diag}(p_i^{-1/2})\sigma ^{1/4}=\sigma ^{-1/4}\), proving the first assertion. The two modular congruences are invertible linear maps, so pre- and post-composition by them preserve the range dimension.

Under the hypotheses of Theorem 13.12.1.1, the matrix

\begin{align} W=(\sigma ^T\otimes \tau )^{-1/4}\rho _{AB} (\sigma ^T\otimes \tau )^{-1/4} \notag \end{align}

is positive semidefinite, and

\begin{align} \operatorname{tr}(W^2)=\| W\| _F^2 & \leq \operatorname {OSR}(\rho _{AB}). \notag \end{align}

This is the faithful-marginal-support form of the order-two estimate obtained from [ Bei13 , Theorem 6, Equation (18) ] .

Proof

The matrix \(W\) is a congruence of \(\rho _{AB}\geq 0\), so it is positive semidefinite and Hermitian; consequently \(\operatorname{tr}(W^2)=\| W\| _F^2\). The weighted map \(L\) is a Hilbert–Schmidt contraction. Its Choi matrix is the whitened state by the preceding theorem, while invertibility of the modular congruences and the supported-marginal correspondence give

\begin{align} \dim _{\mathbb {C}}\operatorname{range}L & =\dim _{\mathbb {C}}\operatorname{range}\Phi _\rho =\operatorname {OSR}(\rho _{AB}). \notag \end{align}

The rectangular Choi range-dimension estimate now proves the claim.

Theorem 13.12.1.3 Real powers under unitary conjugation
#

Let \(A\succeq 0\), let \(s\in \mathbb {R}\), and let \(U\) be unitary. Then

\begin{align} (UAU^\dagger )^s & =UA^sU^\dagger . \notag \end{align}
Proof

Apply Lemma 13.4.22 to the function \(x\mapsto x^s\).

Theorem 13.12.1.4 Real powers of positive tensor products
#

Let \(A\succeq 0\) and \(B\succeq 0\). For every \(s\in \mathbb {R}\),

\begin{align} (A\otimes B)^s & =A^s\otimes B^s. \notag \end{align}
Proof

Apply Theorem 13.4.42 to \(x\mapsto x^s\), using \((xy)^s=x^sy^s\) for \(x,y\geq 0\).

Theorem 13.12.1.5 Inverse quarter-power from iterated square roots

If \(A\succ 0\), then

\begin{align} A^{-1/4} & =\left(\sqrt{\sqrt A}\right)^{-1}. \notag \end{align}
Proof

The continuous functional calculus gives \(\sqrt{\sqrt A}=A^{1/4}\). Positive definiteness makes this matrix invertible, and inversion changes the exponent from \(1/4\) to \(-1/4\).

Theorem 13.12.1.6 Unitary invariance of the order-two functional

Let \(\omega \succeq 0\), let \(\rho \) be a square matrix of the same size, and let \(U\) be unitary. Then

\begin{align} Q_2(U\rho U^\dagger ,U\omega U^\dagger ) & =Q_2(\rho ,\omega ). \notag \end{align}
Proof

Theorem 13.12.1.3 gives \((U\omega U^\dagger )^{-1/4}=U\omega ^{-1/4}U^\dagger \). Substitute this identity into both sandwiching factors and use \(U^\dagger U=1\) and cyclicity of the trace.

Theorem 13.12.1.7 Reindexing invariance of the order-two functional

Let \(\rho \) be a square matrix and let \(\omega \succeq 0\) have the same finite index set. Let \(e\) be a bijection onto another finite index set. Then

\begin{align} Q_2(\rho _{e^{-1},e^{-1}},\omega _{e^{-1},e^{-1}}) & =Q_2(\rho ,\omega ). \notag \end{align}
Proof

Reindexing commutes with the continuous functional calculus, matrix multiplication, and the trace. Apply these three identities to the two inverse quarter-powers and the two copies of the sandwiched matrix.

Let \(\rho _{AB}\succeq 0\) have positive definite marginals \(\rho _A\) and \(\rho _B\). Then

\begin{align} Q_2(\rho _{AB},\rho _A\otimes \rho _B) & \leq \operatorname {OSR}(\rho _{AB}). \label{eq:mpdo_faithful_q2_osr} \end{align}

No trace normalization is required.

Proof

Choose a unitary \(U\) that diagonalizes the first marginal and conjugate \(\rho _{AB}\) by \(U^\dagger \otimes 1\). The two marginals become \(\operatorname{diag}(p_i)\) and \(\rho _B\), while positivity and operator-Schmidt rank are unchanged. The transpose in the Choi whitening convention becomes invisible only in this eigenbasis, since \(\operatorname{diag}(p_i)^T=\operatorname{diag}(p_i)\).

Theorems 13.12.1.4 and 13.12.1.5 identify the negative quarter-power of the product marginal with the two whitening factors in Theorem 13.12.1.2. That theorem bounds the squared whitened trace by the ordinary operator-Schmidt rank. Finally, Theorem 13.12.1.6 and the invariance of operator-Schmidt rank under local unitaries return to the original basis.

Let \(\rho _{AB}\succeq 0\) be any finite-dimensional bipartite operator. Then

\begin{align} Q_2(\rho _{AB},\rho _A\otimes \rho _B) & \leq \operatorname {OSR}(\rho _{AB}). \label{eq:mpdo_support_q2_osr} \end{align}

No trace normalization or positive-dimension hypothesis is required. The statement includes the cases in which either marginal support has dimension zero.

Proof

Compress both factors simultaneously to the supports of \(\rho _A\) and \(\rho _B\). The compressed marginals are positive definite by Theorem 13.11.1.7, even when a support coordinate space has dimension zero. The order-two functional is unchanged by Theorem 13.12.15, and the ordinary operator-Schmidt rank is unchanged by Theorem 13.11.1.5. Apply Theorem 13.12.1.8 to the compressed operator.

The sandwiched Rényi comparison also uses the following unnormalized trace term. For trace-one \(\rho \) and \(\alpha {\gt}1\), it is the quantity inside the logarithm of the sandwiched Rényi divergence [ Bei13 , Equation (3) ] [ MLDS\(^{+}\)13 , Definition 2 ] whenever \(\rho ,\omega \geq 0\) and \(\ker \omega \subseteq \ker \rho \). This includes singular \(\omega \), with negative powers interpreted as the generalized inverse on its support. In this regime, finiteness of the divergence requires the support inclusion. The present application uses \(1{\lt}\alpha \leq 2\); the order-one case is the Umegaki endpoint. The statements below concern only this algebraic expression. They are not assertions of [ CPGSV16 , Proposition 4.5 ] , and they do not establish continuity or monotonicity in the order, interpolation between the endpoints, differentiability at order one, a limit to Umegaki relative entropy, a logarithmic divergence, or a comparison with relative entropy.

Definition 13.12.1.10 Total sandwiched Rényi trace
#

For square complex matrices \(\rho \) and \(\omega \) and every \(\alpha \in \mathbb {R}\), set

\begin{align} \widetilde Q_\alpha (\rho \Vert \omega ) & =\operatorname{Re}\operatorname{tr}\! \left[ \left(\omega ^{(1-\alpha )/(2\alpha )} \rho \, \omega ^{(1-\alpha )/(2\alpha )}\right)^\alpha \right]. \label{eq:sandwiched_renyi_trace} \end{align}

Real powers are defined by continuous functional calculus, with negative powers equal to zero on the zero eigenspace. Thus \(\widetilde Q_\alpha \) is total even when \(\alpha =0\) or \(\omega \) is singular. For \(\alpha {\gt}1\), on positive semidefinite inputs satisfying \(\ker \omega \subseteq \ker \rho \), this convention gives the finite sandwiched Rényi trace term. In this regime, if the support inclusion fails, the totalized value remains finite but is not the divergence trace term.

Theorem 13.12.1.11 Nonnegativity on the faithful domain

If \(\rho \geq 0\) and \(\omega {\gt}0\), then, for every \(\alpha \in \mathbb {R}\),

\begin{align} 0 & \leq \widetilde Q_\alpha (\rho \Vert \omega ). \notag \end{align}
Proof

Put \(q=\omega ^{(1-\alpha )/(2\alpha )}\). Continuous functional calculus gives \(q\geq 0\), hence \(q\rho q\geq 0\) and \((q\rho q)^\alpha \geq 0\). The trace of the last matrix is real and nonnegative.

Theorem 13.12.1.12 Order-one sandwiched trace

If \(\rho \geq 0\) and \(\omega {\gt}0\), then

\begin{align} \widetilde Q_1(\rho \Vert \omega ) & =\operatorname{Re}\operatorname{tr}(\rho ). \notag \end{align}
Proof

At \(\alpha =1\) the two powers of \(\omega \) have exponent zero, while the remaining real power of \(\rho \) has exponent one. Substitution into (212) gives the identity.

If \(\rho \geq 0\) and \(\omega {\gt}0\), then

\begin{align} \widetilde Q_2(\rho \Vert \omega ) & =Q_2(\rho ,\omega ). \notag \end{align}
Proof

At \(\alpha =2\) the sandwiching exponent is \(-1/4\). Since \(\omega ^{-1/4}\rho \, \omega ^{-1/4}\) is positive semidefinite, its continuous-functional-calculus square is its ordinary matrix square.

Theorem 13.12.1.14 Mutual information from operator-Schmidt rank

Let \(\rho _{AB}\succeq 0\) be a finite-dimensional bipartite operator with \(\operatorname{tr}\rho _{AB}=1\). Then

\begin{align} I(A:B)_\rho \leq \log \operatorname {OSR}(\rho _{AB}). \notag \end{align}

No positive-dimension or faithful-marginal hypothesis is required. The operator-Schmidt rank is the ordinary operator-Schmidt rank.

Proof

Set \(\omega =\rho _A\otimes \rho _B\). The marginals are positive semidefinite, \(\operatorname{tr}\omega =1\), and Theorem 13.11.1.4 gives \(\ker \omega \subseteq \ker \rho _{AB}\). Reindex the product basis by a single finite coordinate set. Lemma 13.4.33 and Theorem 13.12.1.7 preserve the two functionals, so Theorem 13.12.17 applies. Theorem 13.6.46 identifies its left side with mutual information. Thus, with \(q=Q_2(\rho _{AB},\omega )\),

\begin{align} I(A:B)_\rho & =D(\rho _{AB}\Vert \omega ) \leq \log q, \notag \\ 0 & \leq q\leq \operatorname {OSR}(\rho _{AB}). \notag \end{align}

The second inequality is Theorem 13.12.1.9. Since \(\operatorname{tr}\rho _{AB}=1\), the operator is nonzero, and Theorem 13.11.1.1 gives \(\operatorname {OSR}(\rho _{AB}){\gt}0\). This positive integer is at least one, so its logarithm is nonnegative. If \(q=0\), the totalized convention \(\log 0=0\) gives \(I(A:B)_\rho \leq 0\leq \log \operatorname {OSR}(\rho _{AB})\). If \(q{\gt}0\), monotonicity of the logarithm gives \(\log q\leq \log \operatorname {OSR}(\rho _{AB})\).

13.13 Trivial-factor corollaries

The partial traces and von Neumann entropy are introduced in Chapter 13. This supplement collects elementary trace-preservation identities and the dimension-1 entropy bound that follow directly from those definitions. They are listed as separate results because each proof requires only unfolding definitions and re-indexing finite sums.

Theorem 13.13.1 Partial trace over \(A\) preserves the full trace
#

For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_A(\rho _{ABC}))\).

Proof

Unfolding the definitions,

\begin{align} \operatorname{tr}(\operatorname{tr}_A(\rho _{ABC})) & = \sum _{b,c}\sum _a (\rho _{ABC})_{(a,b,c)(a,b,c)} = \operatorname{tr}(\rho _{ABC}). \notag \end{align}
Theorem 13.13.2 Partial trace over \(C\) preserves the full trace
#

For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_C(\rho _{ABC}))\).

Proof

Unfolding the definitions,

\begin{align} \operatorname{tr}(\operatorname{tr}_C(\rho _{ABC})) & = \sum _{a,b}\sum _c (\rho _{ABC})_{(a,b,c)(a,b,c)} = \operatorname{tr}(\rho _{ABC}). \notag \end{align}
Theorem 13.13.3 Partial trace over \(AC\) preserves the full trace
#

For any tripartite matrix \(\rho _{ABC}\), \(\operatorname{tr}(\rho _{ABC})=\operatorname{tr}(\operatorname{tr}_{AC}(\rho _{ABC}))\).

Proof

Unfolding the definitions and interchanging the outer sums,

\begin{align} \operatorname{tr}(\operatorname{tr}_{AC}(\rho _{ABC})) & = \sum _b\sum _{a,c} (\rho _{ABC})_{(a,b,c)(a,b,c)} = \operatorname{tr}(\rho _{ABC}). \notag \end{align}
Theorem 13.13.4 Entropy vanishes in dimension \(1\)

If \(\rho \in M_{1}(\mathbb {C})\) is Hermitian and \(\operatorname{tr}(\rho )=1\), then \(S(\rho )=0\).

Proof

The unique eigenvalue equals the trace, hence equals \(1\), and \(-1\cdot \log 1=0\):

\begin{align} \lambda _0 & = \operatorname{tr}(\rho )=1, \notag \\ S(\rho ) & = -\lambda _0\log \lambda _0 =-1\cdot \log 1=0. \notag \end{align}