Tensor Network Theory: A formalization blueprint

20 Operator Convexity and Jensen Inequalities

This chapter collects the operator convexity/concavity results, trace inequalities, the operator Jensen inequality for positive maps, and the Lieb concavity theorem. These are standard results in matrix analysis ( [ Bha97 , Chapter V ] ; [ Wol12 , Theorems 5.12 and 5.13 ] ; Hansen–Pedersen’s Jensen inequality  [ HP82 ] ; and Lieb’s concavity theorem  [ Lie73 ] ). The operator Jensen inequalities for real powers \(x\mapsto x^p\) on the Loewner order are obtained from the Loewner integral representation of the power function. Lieb’s concavity theorem is derived from the operator integral representation of the fractional product together with the vectorization isometry, which turns the trace functional into a quadratic form of a positive-semidefinite matrix.

20.1 Schwarz inequalities for positive maps

The following consequences of positivity and the Kadison–Schwarz inequality concern normal, subnormal, and commuting-dominant operators. They provide matrix-order tools used independently of the Fundamental Theorem.

20.1.1 Schwarz inequality for normal operators

Theorem 20.1.1.1 CP Schwarz inequality for normal operators

Let \(E^*(X)=\sum _i K_i^\dagger X K_i\) be the adjoint Kraus map of a trace-preserving Kraus family, and let \(A\in M_{D}(\mathbb {C})\) be normal. Then the Schwarz gap

\begin{align} E^*(A^\dagger A)-E^*(A^\dagger )E^*(A) \label{eq:schwarz_normal_cp_gap} \end{align}

is positive semidefinite. Equivalently,

\begin{align} E^*(A^\dagger )E^*(A) \le E^*(A^\dagger A). \label{eq:schwarz_normal_cp_loewner} \end{align}

This is the completely positive case of [ Wol12 , Proposition 5.1 ] .

Proof

Positivity of the gap in (??) is exactly Theorem 5.1.6. The Loewner inequality (??) is its order-theoretic reformulation.

Theorem 20.1.1.2 Schwarz inequality for normal operators

Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be a positive linear map with \(T(\mathbb {1})\le \mathbb {1}\). If \(A\in M_{D}(\mathbb {C})\) is normal, then \(T(A^\dagger A)-T(A^\dagger )T(A)\ge 0\). This is [ Wol12 , Proposition 5.1 ] .

Proof

Since \(A\) is normal, diagonalize \(A=UDU^\dagger \) with \(D\) diagonal and \(U\) unitary. The spectral projectors of \(A\) generate a commutative \(*\)-subalgebra. The restriction of \(T\) to this subalgebra is completely positive, so the Kadison–Schwarz inequality for completely positive subunital maps gives \(T(A^\dagger )T(A)\le T(A^\dagger A)\).

Remark 20.1.1.3
#

The proof follows the positivity-on-abelian argument of [ Wol12 , Proposition 1.6 ] : a positive map restricted to the commutative \(*\)-subalgebra generated by the spectral projectors of \(A\) satisfies the Schwarz inequality directly, without invoking the Arveson extension theorem.

20.1.2 Schwarz inequality for subnormal and commuting-dominant operators

Theorem 20.1.2.1 Schwarz inequality for subnormal operators

Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be a positive linear map with \(T(\mathbb {1})\le \mathbb {1}\). If \(A\in M_{D}(\mathbb {C})\) is subnormal (i.e. there exists a normal operator on a larger space whose compression to \(\mathbb {C}^D\) is \(A\)), then \(T(A^\dagger A)-T(A^\dagger )T(A)\ge 0\). This is [ Wol12 , Theorem 5.5 ] .

Proof

Embed \(A\) as the northwest corner of a normal operator \(N\) on \(\mathbb {C}^D\oplus \mathbb {C}^E\). Extend \(T\) to a positive subunital map \(\widetilde{T}\) on the larger space, acting as \(T\) on the northwest block and as zero elsewhere. Apply Theorem 20.1.1.2 to \(\widetilde{T}\) and \(N\), then compress the resulting inequality back to \(\mathbb {C}^D\).

Theorem 20.1.2.2 Schwarz inequality for commuting-dominant operators

Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be a positive linear map with \(T(\mathbb {1})\le \mathbb {1}\). If \(A\in M_{D}(\mathbb {C})\) admits a positive semidefinite dominant \(B\ge 0\) with \(A^\dagger A\le B\) and \(B\) commuting with \(A\) (\(BA=AB\)), then

\begin{align} T(A^\dagger )T(A)& \le T(B), \notag \\ T(A)T(A^\dagger )& \le T(B). \notag \end{align}

This is [ Wol12 , Theorem 5.6 ] .

Proof

First take \(B{\gt}0\). Since \(B\) commutes with \(A\) and \(A^\dagger A\le B\), also \(AA^\dagger \le B\), so \(D_R=(B-A^\dagger A)^{1/2}\) and \(D_L=(B-AA^\dagger )^{1/2}\) are positive semidefinite. The block matrix

\begin{align} N & = \begin{pmatrix} A & D_L \\ D_R & -A^\dagger \end{pmatrix}, \notag \\ N^\dagger N & =NN^\dagger =B\oplus B. \notag \end{align}

is normal. As in the subnormal case, extend \(T\) to a positive subunital map on \(M_{2D}(\mathbb {C})\), apply the normal-operator Schwarz inequality (Theorem 20.1.1.2) to that extension and \(N\), and compress back to \(\mathbb {C}^D\). Since \(N\) is normal, both \(T(A^\dagger )T(A)\le T(B)\) and \(T(A)T(A^\dagger )\le T(B)\) follow. For general \(B\ge 0\), apply this to \(B_\varepsilon =B+\varepsilon \mathbb {1}{\gt}0\) and let \(\varepsilon \to 0\).

Theorem 20.1.2.3 CP variant: Kadison–Schwarz for commuting-dominant inputs

Let \(E^*\) be the adjoint of a trace-preserving Kraus map. If \(A\in M_{D}(\mathbb {C})\) admits a positive semidefinite dominant \(B\) commuting with \(A\) and satisfying \(A^\dagger A\le B\), then

\begin{align} E^*(A^\dagger )E^*(A)& \le E^*(B), \notag \\ E^*(A)E^*(A^\dagger )& \le E^*(B). \notag \end{align}
Proof

Apply Theorem 20.1.2.2 to the adjoint Kraus map, which is unital under the trace-preservation hypothesis.

20.2 Diagonal Jensen inequality

Lemma 20.2.1 Diagonal Jensen inequality
#

Let \(f:\mathbb {R}\to \mathbb {R}\) be convex on \([0,\infty )\), let \(A\) be a positive semidefinite matrix on a finite-dimensional space indexed by \(n\), and let \(v\in \mathbb {C}^n\) satisfy \(\langle v,v\rangle =1\). Then

\begin{align} f\! \left(\operatorname{Re}\langle v,Av\rangle \right) \le \operatorname{Re}\langle v,f(A)v\rangle , \notag \end{align}

where \(f(A)\) is defined by the Hermitian continuous functional calculus.

Proof

Write \(A=U\operatorname{diag}(\mu )U^\dagger \) by the spectral theorem and set \(w=U^\dagger v\). Since \(U\) is unitary, \(p_i:=\lvert w_i\rvert ^2\) satisfies \(\sum _i p_i=\lVert v\rVert ^2=1\), so \((p_i)\) is a probability distribution on \(n\). The eigenvalues \(\mu _i\) lie in \([0,\infty )\) because \(A\) is positive semidefinite. A direct computation yields

\begin{align} \operatorname{Re}\langle v,Av\rangle & =\sum _i p_i\mu _i, \notag \\ \operatorname{Re}\langle v,f(A)v\rangle & =\sum _i p_i f(\mu _i). \notag \end{align}

so the scalar Jensen inequality applied to the weights \((p_i)\) and points \((\mu _i)\) gives the conclusion. See [ Bha97 , Chapter V ] .

20.3 Trace concavity and convexity of matrix powers

Theorem 20.3.1 Trace concavity of real powers

For \(p\in [0,1]\), PSD matrices \(A_1,A_2\), and \(t\in [0,1]\),

\begin{align} t\, \operatorname{Re}\operatorname{tr}(A_1^p)+(1-t)\operatorname{Re}\operatorname{tr}(A_2^p) \le \operatorname{Re}\operatorname{tr}\! \left((tA_1+(1-t)A_2)^p\right). \notag \end{align}

This follows from operator concavity of \(x\mapsto x^p\) composed with trace monotonicity on the Loewner order.

Proof

Let \(A=tA_1+(1-t)A_2\), and let \(\{ \psi _j\} \) be an orthonormal eigenbasis of \(A\) with eigenvalues \(\mu _j\ge 0\). A direct computation gives

\begin{align} \operatorname{Re}\operatorname{tr}(A^p) & =\sum _j\mu _j^p, \notag \\ \mu _j & =t\, a_j+(1-t)b_j, \notag \\ a_j & =\operatorname{Re}\langle \psi _j,A_1\psi _j\rangle , \notag \\ b_j & =\operatorname{Re}\langle \psi _j,A_2\psi _j\rangle . \notag \end{align}

Scalar concavity of \(x\mapsto x^p\) on \([0,\infty )\) yields \(\mu _j^p\ge t\, a_j^p+(1-t)b_j^p\). Applying Lemma 20.2.1 to the convex function \(x\mapsto -x^p\) and multiplying by \(-1\) gives

\begin{align} a_j^p & \ge \operatorname{Re}\langle \psi _j,A_1^p\psi _j\rangle , \notag \\ b_j^p & \ge \operatorname{Re}\langle \psi _j,A_2^p\psi _j\rangle . \notag \end{align}

Summing over \(j\) and using \(\sum _j\operatorname{Re}\langle \psi _j,M\psi _j\rangle =\operatorname{Re}\operatorname{tr}M\) yields the conclusion. See [ Bha97 , Chapter V ] .

Theorem 20.3.2 Trace convexity of real powers

For \(p\in [1,2]\), PSD matrices \(A_1,A_2\), and \(t\in [0,1]\),

\begin{align} \operatorname{Re}\operatorname{tr}\! \left((tA_1+(1-t)A_2)^p\right) \le t\, \operatorname{Re}\operatorname{tr}(A_1^p)+(1-t)\operatorname{Re}\operatorname{tr}(A_2^p). \notag \end{align}

This follows from operator convexity of \(x\mapsto x^p\) composed with trace monotonicity on the Loewner order.

Proof

Use the eigenbasis reduction from Theorem 20.3.1, with scalar convexity rather than concavity of \(x\mapsto x^p\) on \([0,\infty )\) for \(p\ge 1\) and the convex form of the diagonal Jensen inequality.

20.4 Operator convexity of real powers

Theorem 20.4.1 Operator convexity of real powers between one and two
#

For elements \(a,b\) of a unital C\({}^\ast \)-algebra with \(0\le a,b\), an exponent \(p\in [1,2]\), and \(t\in [0,1]\), the map \(a\mapsto a^p\) is convex for the Loewner order on the positive cone:

\begin{align} (ta+(1-t)b)^p \le t\, a^p+(1-t)b^p. \notag \end{align}

This is the convex counterpart of the operator concavity of \(x\mapsto x^p\) for exponents \(p\in [0,1]\).

Proof

For \(p\in (1,2)\) one uses the integral representation \(a^p=\int _0^\infty f_{p,s}(a)\, d\mu (s)\) of Carlen [ Car10 , Lemma 2.8 ] , whose integrand \(f_{p,s}(x)=s^{p-1}(s^{-1}x+s(s+x)^{-1}-1)\) is operator convex: it decomposes into a linear term, a non-negative multiple of the operator convex resolvent \(x\mapsto (s\cdot 1+x)^{-1}\), and a constant. Since \(f_{p,s}(ta+(1-t)b)\le t\, f_{p,s}(a)+(1-t)f_{p,s}(b)\) for every \(s{\gt}0\), integrating this pointwise Loewner inequality against \(\mu \) gives

\begin{align} (ta+(1-t)b)^p & =\int _0^\infty f_{p,s}(ta+(1-t)b)\, d\mu (s) \notag \\ & \le t\int _0^\infty f_{p,s}(a)\, d\mu (s) +(1-t)\int _0^\infty f_{p,s}(b)\, d\mu (s) \notag \\ & =t\, a^p+(1-t)b^p. \notag \end{align}

The endpoint \(p=1\) is the identity, and \(p=2\) is \(a\mapsto a^2\), operator convex because \(t\, a^2+(1-t)b^2-(ta+(1-t)b)^2=t(1-t)(a-b)^2\ge 0\). See [ Bha97 , Chapter V ] .

20.5 Operator Jensen inequality for positive maps

Definition 20.5.1 Finite-POVM dilation construction

For a finite family of matrices \(\{ C_i\} _{i\in \iota }\) and a defect block \(S\), one forms an isometric dilation whose rows are the adjoints of the matrices \(C_i\) together with the adjoint of \(S\). One also forms the scalar block-diagonal matrix whose \(\iota \)-blocks carry prescribed weights and whose defect block carries a single scalar.

Lemma 20.5.2 Compression auxiliary for Jensen
#

Let \(Y\) be positive definite and let \(W\) be an isometry. Then

\begin{align} (W^\dagger YW)^{-1} \le W^\dagger Y^{-1}W. \label{eq:opconv_inverse_compression_le} \end{align}
Proof

Apply the Schur-complement criterion to the block matrix \(\begin{psmallmatrix} \end{psmallmatrix}Y^{-1}& W\\ W^\dagger & W^\dagger YW\end{psmallmatrix}\). The lower-right Schur complement is zero, hence positive semidefinite, and compressing the resulting upper-left Schur complement by \(W\) gives (??).

If the defect block satisfies

\begin{align} SS^\dagger =\mathbb {1}-\sum _{i\in \iota }C_iC_i^\dagger , \label{eq:opconv_defect_block_relation} \end{align}

then the dilation is an isometry, and compressing the weighted scalar block-diagonal matrix gives \(\sum _{i\in \iota }w_iC_iC_i^\dagger +tSS^\dagger \). The defect relation can also be written as \(\sum _{i\in \iota }C_iC_i^\dagger +SS^\dagger =\mathbb {1}\). Strict positivity of all weights together with \(t{\gt}0\) implies positive definiteness of the scalar block-diagonal matrix. If all weights are nonzero and the defect scalar is nonzero, the inverse diagonal has reciprocal weights, and compressing this inverse gives \(\sum _{i\in \iota }w_i^{-1}C_iC_i^\dagger +t^{-1}SS^\dagger \).

Proof

Expand the block matrix multiplication entrywise. The defect identity (??) gives the isometry relation after summing the \(\iota \)-blocks and the defect block, and the weighted compression identity follows from the same expansion with the scalar diagonal inserted. Positive definiteness uses strict positivity of every diagonal entry, while the inverse formula uses the nonvanishing of every diagonal entry.

Lemma 20.5.4 Finite-POVM resolvent inequality

Let \(w_i\ge 0\) and \(t{\gt}0\). If the defect block satisfies \(SS^\dagger =\mathbb {1}-\sum _{i\in \iota }C_iC_i^\dagger \), then

\begin{align} \left(\sum _{i\in \iota }w_iC_iC_i^\dagger +t\mathbb {1}\right)^{-1} \le \sum _{i\in \iota }(w_i+t)^{-1}C_iC_i^\dagger +t^{-1}SS^\dagger . \label{eq:opconv_povm_resolvent_bound} \end{align}
Proof

Apply the compression inequality (??) to the positive definite scalar block-diagonal matrix with entries \(w_i+t\) and defect entry \(t\). The compression identities rewrite the compressed matrix as \(\sum _iw_iC_iC_i^\dagger +t\mathbb {1}\) and its inverse compression as the upper bound in (??).

Lemma 20.5.5 Spectral finite-POVM resolution

Let \(A\ge 0\) be a finite matrix. There are positive semidefinite matrices \(P_i\) and non-negative real numbers \(\lambda _i\) such that

\begin{align} \sum _iP_i& =\mathbb {1}, \notag \\ A& =\sum _i\lambda _iP_i. \notag \end{align}
Proof

Diagonalize the Hermitian matrix \(A\) by a unitary matrix. The matrices \(P_i\) are the conjugates of the rank-one diagonal matrix units, their sum is the identity, and positive semidefiniteness of \(A\) gives \(\lambda _i\ge 0\).

Lemma 20.5.6 Positive-map resolvent inequality

Let \(T\) be a positive subunital map, let \(A\ge 0\), and let \(t{\gt}0\). Then

\begin{align} (T(A)+t\mathbb {1})^{-1} \le T\! \left((A+t\mathbb {1})^{-1}\right)+t^{-1}(\mathbb {1}-T(\mathbb {1})). \notag \end{align}
Proof

Use the spectral resolution \(A=\sum _i\lambda _iP_i\). The matrices \(T(P_i)\) are positive semidefinite and sum to \(T(\mathbb {1})\le \mathbb {1}\). Write each \(T(P_i)\) as \(C_iC_i^\dagger \) using the positive square root, and write the defect \(\mathbb {1}-\sum _iT(P_i)\) as \(SS^\dagger \). The finite-POVM resolvent inequality gives the desired estimate after substituting the spectral formula for \((A+t\mathbb {1})^{-1}\).

Lemma 20.5.7 Positive-map real-power integrand inequality

Let \(T\) be a positive subunital map, let \(A\ge 0\), let \(p\in (0,1)\), and let \(t{\gt}0\). Then

\begin{align} T\! \left(f_{p,t}(A)\right) & \le f_{p,t}(T(A)), \notag \\ f_{p,t}(x) & =t^p\left(t^{-1}-(t+x)^{-1}\right). \notag \end{align}
Proof

Rewrite the continuous functional calculus in resolvent form as \(f_{p,t}(X)=t^{p-1}\mathbb {1}-t^p(t\mathbb {1}+X)^{-1}\). After applying the linear map \(T\), the difference \(f_{p,t}(T(A))-T(f_{p,t}(A))\) is exactly \(t^p\) times the positive matrix supplied by the positive-map resolvent inequality.

Lemma 20.5.8 Positive-map convex real-power integrand inequality

Let \(T\) be a positive subunital map, let \(A\ge 0\), let \(p\in (1,2)\), and let \(t{\gt}0\). Then

\begin{align} g_{p,t}(T(A)) & \le T\! \left(g_{p,t}(A)\right), \notag \\ g_{p,t}(x) & =t^{p-2}x+t^p(t+x)^{-1}-t^{p-1}. \notag \end{align}

The inequality is reversed relative to the concave integrand, since \(g_{p,t}\) is operator convex.

Proof

Rewrite the continuous functional calculus in resolvent form as \(g_{p,t}(X)=t^{p-2}X+t^p(t\mathbb {1}+X)^{-1}-t^{p-1}\mathbb {1}\). Under the linear map \(T\), the leading \(t^{p-2}\) terms cancel and the constant \(t^{p-1}\) terms combine, so the difference \(T(g_{p,t}(A))-g_{p,t}(T(A))\) is exactly \(t^p\) times the positive matrix supplied by the positive-map resolvent inequality, using \(t^pt^{-1}=t^{p-1}\).

Lemma 20.5.9 Positive matrix-valued integrals

Let \(f\) be a function from a measure space to a finite matrix algebra. If \(f(x)\ge 0\) for almost every \(x\), then \(\int f(x)\, d\mu (x)\ge 0\).

Proof

The Loewner order has closed positive cone in finite dimension. Hence the ordered Bochner integral preserves almost-everywhere nonnegativity. The corresponding almost-everywhere monotonicity statement is the general ordered Bochner integral theorem.

Theorem 20.5.10 Jensen for concave real powers

For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [0,1]\), one has \(T(A^p)\le (T(A))^p\). This is [ Wol12 , Theorem 5.13 ] for concave \(f(x)=x^p\).

Proof

For \(p=0\) the assertion is the subunitality condition \(T(\mathbb {1})\le \mathbb {1}\). For \(p=1\) it is an equality. For \(0{\lt}p{\lt}1\), use Carlen’s integral representation of \(A^p\) by the functions \(f_{p,t}(x)=t^p(t^{-1}-(t+x)^{-1})\). The pointwise inequality \(T(f_{p,t}(A))\le f_{p,t}(T(A))\) is Lemma 20.5.7. Integrating this inequality with respect to the representing measure gives the result by Lemma 20.5.9 and linearity of the Bochner integral under \(T\).

Theorem 20.5.11 Map form of Jensen for concave real powers
#

For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [0,1]\), the inequality \(T(A^p)\le (T(A))^p\) holds.

Proof

Direct application of Theorem 20.5.10.

Theorem 20.5.12 Jensen for convex real powers

For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [1,2]\), one has \((T(A))^p\le T(A^p)\). This is [ Wol12 , Theorem 5.11 ] for the subunital operator-convex case \(f(x)=x^p\).

Proof

For \(p=1\) the assertion is an equality. For \(1{\lt}p{\lt}2\), use Carlen’s integral representation of \(A^p\) by the convex integrands \(g_{p,t}(x)=t^{p-2}x+t^p(t+x)^{-1}-t^{p-1}\). The pointwise inequality \(g_{p,t}(T(A))\le T(g_{p,t}(A))\) is Lemma 20.5.8. Integrating this inequality with respect to the representing measure and using Lemma 20.5.9 together with linearity of the Bochner integral under \(T\) gives the inequality for \(1{\lt}p{\lt}2\). The endpoint \(p=2\) follows by taking the limit \(q\to 2^-\) in \((T(A))^q\le T(A^q)\): since \(M^q\to M^2\) in the operator norm as \(q\to 2^-\) for any positive semidefinite \(M\), and the Loewner order is closed, this gives \((T(A))^2\le T(A^2)\).

Theorem 20.5.13 Map form of Jensen for convex real powers
#

For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [1,2]\), the inequality \((T(A))^p\le T(A^p)\) holds.

Proof

Direct application of Theorem 20.5.12.

Theorem 20.5.14 Jensen for concave log

For a positive unital map \(T\) and positive-definite \(A\), one has \(T(\log A)\le \log (T(A))\). This is [ Wol12 , Theorem 5.13 ] for \(f=\log \). Requires unitality (\(T(\mathbb {1})=\mathbb {1}\)), not merely subunitality.

Proof

Apply Theorem 20.5.10 to \(0{\lt}p{\lt}1\) and rewrite it as an inequality for \((A^p-\mathbb {1})/p\). The continuous functional calculus gives \((A^p-\mathbb {1})/p\to \log A\) as \(p\to 0^+\), and the same limit holds for \(T(A)\). Since the Loewner order graph is closed in finite dimension, the limiting inequality is \(T(\log A)\le \log (T(A))\).

Theorem 20.5.15 Map form of Jensen for the concave logarithm
#

For a positive unital map \(T\) and a positive-definite matrix \(A\), the inequality \(T(\log A)\le \log (T(A))\) holds.

Proof

Direct application of Theorem 20.5.14.

20.6 Lieb concavity theorem

The Lieb concavity theorem rests on a scalar integral identity that expresses the geometric interpolation \(a^s b^{1-s}\) as a positive average of the elementary fractions \(ab/(a+tb)\). This is the analytic core of the integral representation of \(A^s B^{1-s}\).

Lemma 20.6.1 Reflection integral
#

For \(s\in (0,1)\),

\begin{align} \int _0^\infty \frac{u^{s-1}}{1+u}\, du & = \frac{\pi }{\sin (\pi s)}. \notag \end{align}
Proof

This is the value \(B(s,1-s)\) of the Beta function, which equals \(\Gamma (s)\Gamma (1-s)=\pi /\sin (\pi s)\) by Euler’s reflection formula. The substitution \(x=u/(1+u)\) carries the half-line onto the unit interval and turns the integrand into the Beta integrand \(x^{s-1}(1-x)^{-s}\).

Lemma 20.6.2 Integrability of the Lieb integrand
#

For \(a,b{\gt}0\) and \(s\in (0,1)\), the function \(t\mapsto t^{s-1} \dfrac {ab}{a+tb}\) is integrable on \((0,\infty )\).

Proof

The integrand is \(O(t^{s-1})\) as \(t\to 0^+\) and \(O(t^{s-2})\) as \(t\to \infty \), so it has integrable comparison functions on each tail.

Lemma 20.6.3 Value of the Lieb integral
#

For \(a,b{\gt}0\) and \(s\in (0,1)\),

\begin{align} \int _0^\infty t^{s-1} \frac{ab}{a+tb}\, dt & = a^s b^{1-s} \frac{\pi }{\sin (\pi s)}. \notag \end{align}
Proof

The scaling substitution \(t=(a/b)\, u\) turns the integrand into \(a^s b^{1-s}\, u^{s-1}/(1+u)\), whose half-line integral is the reflection integral \(\pi /\sin (\pi s)\) of Lemma 20.6.1.

Lemma 20.6.4 Scalar Lieb integral identity
#

For \(a,b{\gt}0\) and \(s\in (0,1)\),

\begin{align} a^s b^{1-s} & = \frac{\sin (\pi s)}{\pi } \int _0^\infty t^{s-1} \frac{ab}{a+tb}\, dt. \notag \end{align}
Proof

By Lemma 20.6.3, the integral equals \(a^s b^{1-s} \pi /\sin (\pi s)\), so

\begin{align} \frac{\sin (\pi s)}{\pi } \int _0^\infty t^{s-1} \frac{ab}{a+tb}\, dt & = \frac{\sin (\pi s)}{\pi }\cdot a^s b^{1-s} \frac{\pi }{\sin (\pi s)} = a^s b^{1-s}. \notag \end{align}
Lemma 20.6.5 Diagonal Lieb integral identity

For positive diagonal matrices \(M\) and \(N\) and \(s\in (0,1)\),

\begin{align} M^s N^{1-s} & = \frac{\sin (\pi s)}{\pi } \int _0^\infty t^{s-1}\, M\, (M+tN)^{-1}\, N\, dt. \notag \end{align}
Proof

Both sides are diagonal, so the identity holds entry by entry; on each diagonal entry it is the scalar Lieb integral identity (Lemma 20.6.4).

Lemma 20.6.6 Operator integral representation
#

For positive-definite matrices \(A\), \(B\) and \(s\in (0,1)\),

\begin{align} (A\otimes \mathbb {1})^s\, (\mathbb {1}\otimes B^\top )^{1-s} & = \frac{\sin (\pi s)}{\pi } \int _0^\infty t^{s-1}\, (A\otimes \mathbb {1}) ((A\otimes \mathbb {1})+t(\mathbb {1}\otimes B^\top ))^{-1} (\mathbb {1}\otimes B^\top )\, dt, \label{eq:operator_convexity_integral_rep} \end{align}

where \(A\otimes \mathbb {1}\) and \(\mathbb {1}\otimes B^\top \) are regarded as operators on the Kronecker model space \(\mathbb {C}^{D\times D}\).

Proof

The pair \((A\otimes \mathbb {1},\mathbb {1}\otimes B^\top )\) is simultaneously diagonalized by \(V=U_A\otimes U_{B^\top }\) (eigenbases of \(A\) and \(B^\top \)). In that basis the pair becomes a commuting pair of positive diagonal matrices, and the identity follows from the diagonal case (Lemma 20.6.5).

Lemma 20.6.7 Concavity of the parallel sum
#

For positive-definite matrices \(X_1,X_2,Y_1,Y_2\) and \(\theta \in [0,1]\), the parallel sum \((X,Y)\mapsto X(X+Y)^{-1}Y\) is jointly concave in the Loewner order:

\begin{align} \theta \, X_1(X_1+Y_1)^{-1}Y_1 +(1-\theta )\, X_2(X_2+Y_2)^{-1}Y_2 & \le X_\theta (X_\theta +Y_\theta )^{-1}Y_\theta , \notag \end{align}

where \(X_\theta =\theta X_1+(1-\theta )X_2\) and \(Y_\theta =\theta Y_1+(1-\theta )Y_2\).

Proof

Write \(Z=X(X+Y)^{-1}Y\). The identity \(X(X+Y)^{-1}(X+Y)=X\) gives \((X-Z)-X(X+Y)^{-1}X=0\), so the block matrix \(\begin{psmallmatrix} \end{psmallmatrix}X-Z & X\\ X & X+Y\end{psmallmatrix}\) is positive semidefinite (its Schur complement against \(X+Y\) vanishes). Setting \(Z_i=X_i(X_i+Y_i)^{-1}Y_i\) and \(Z_\theta =\theta Z_1+(1-\theta )Z_2\), the convex combination of the two block matrices is \(\begin{psmallmatrix} \end{psmallmatrix}X_\theta -Z_\theta & X_\theta \\ X_\theta & X_\theta +Y_\theta \end{psmallmatrix}\), which is positive semidefinite, and its Schur complement against \(X_\theta +Y_\theta \) yields \(X_\theta (X_\theta +Y_\theta )^{-1}Y_\theta -Z_\theta \ge 0\), which is the claimed inequality (Ando, 1979).

Lemma 20.6.8 Concavity of the resolvent integrand
#

For \(t{\gt}0\), positive-definite \(A_1,A_2,B_1,B_2\), and \(\theta \in [0,1]\), with \(\hat{A}=A\otimes \mathbb {1}\) and \(\hat{B}=\mathbb {1}\otimes B^\top \), setting \(A_\theta =\theta A_1+(1-\theta )A_2\) and \(B_\theta =\theta B_1+(1-\theta )B_2\):

\begin{align} \theta \, \hat A_1(\hat A_1+t\hat B_1)^{-1}\hat B_1 +(1-\theta ) \hat A_2(\hat A_2+t\hat B_2)^{-1}\hat B_2 & \le \hat A_\theta (\hat A_\theta +t\hat B_\theta )^{-1}\hat B_\theta . \notag \end{align}
Proof

The integrand equals \(t^{-1} \hat A(\hat A+t\hat B)^{-1}(t\hat B)\), a rescaled parallel sum of \(\hat A\) and \(t\hat B\), so concavity follows from Lemma 20.6.7 and the linearity of \(A\mapsto \hat A\), \(B\mapsto \hat B\).

Lemma 20.6.9 Concavity of the fractional product
#

For positive-definite matrices \(A_1,A_2,B_1,B_2\), \(s\in (0,1)\), and \(\theta \in [0,1]\), with \(\hat{A}=A\otimes \mathbb {1}\) and \(\hat{B}=\mathbb {1}\otimes B^\top \), setting \(A_\theta =\theta A_1+(1-\theta )A_2\) and \(B_\theta =\theta B_1+(1-\theta )B_2\), the fractional product \(\hat A^s\hat B^{1-s}\) is jointly concave in the Loewner order:

\begin{align} \theta \, \hat A_1^{\, s}\hat B_1^{\, 1-s} +(1-\theta ) \hat A_2^{\, s}\hat B_2^{\, 1-s} & \le \hat A_\theta ^{\, s}\hat B_\theta ^{\, 1-s}. \notag \end{align}
Proof

By the integral representation (??), each fractional product equals the positive multiple \(\frac{\sin (\pi s)}{\pi }\) of the integral of the resolvent integrand against the non-negative weight \(t^{s-1}\) on \((0,\infty )\). For each \(t{\gt}0\), the integrand is jointly concave by Lemma 20.6.8, so the difference of the averaged point and the convex combination is, almost everywhere, the non-negative scalar \(t^{s-1}\) times a positive-semidefinite matrix. Integrating a positive-semidefinite integrand yields a positive-semidefinite matrix, and scaling by the positive constant \(\frac{\sin (\pi s)}{\pi }\) preserves the Loewner order, which gives the inequality.

Theorem 20.6.10 Lieb concavity
#

For \(s\in [0,1]\), any matrix \(K\), and positive-definite matrices \(A_1,A_2,B_1,B_2\), write \(F_s(A,B):=\Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\). This map is jointly concave. For \(t\in [0,1]\), set \(A_t:=tA_1+(1-t)A_2\) and \(B_t:=tB_1+(1-t)B_2\). Then

\begin{align} tF_s(A_1,B_1)+(1-t)F_s(A_2,B_2) & \le F_s(A_t,B_t). \label{eq:opconv_lieb_joint_concavity} \end{align}

See [ Lie73 ; And79 ] . This statement is the positive-definite, boundary-exponent (\(x+y=1\)) case of the Ando–Lieb theorem [ Wol12 , Theorem 5.15 ] , which holds for positive semidefinite \(A,B\) and all exponents \(x,y\ge 0\) with \(x+y\le 1\).

Proof

For \(s\in (0,1)\), the concavity of the fractional product of the commuting left and right multiplication operators (Lemma 20.6.9) gives a Loewner-order inequality between positive-semidefinite matrices on the model space \(\mathbb {C}^D\otimes \mathbb {C}^D\). Writing that product as \(A^s\otimes (B^\top )^{1-s}\) and pairing it with the vectorization of \(K^\top \),

\begin{align} \left\langle \operatorname{vec}K^\top , (A^s\otimes (B^\top )^{1-s})\operatorname{vec}K^\top \right\rangle & = \Re \operatorname{tr}(K^\dagger A^s K B^{1-s}), \notag \end{align}

expresses the trace functional as the quadratic form of the Kronecker product at that vector. Applying this positive quadratic form to the Loewner inequality transfers the concavity to the real part of the trace. The endpoints \(s=0\) and \(s=1\) reduce the varying side to a linear function of the convex-combination argument, hence hold with equality.

Theorem 20.6.11 Lieb concavity, positive-semidefinite inputs
#

For \(s\in [0,1]\), any matrix \(K\), and positive-semidefinite matrices \(A_1,A_2,B_1,B_2\), the map \((A,B)\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is jointly concave, satisfying (??) of Theorem 20.6.10. This is the full Ando–Lieb theorem [ Wol12 , Theorem 5.15 ] on the boundary line \(x+y=1\) (with \(x=s\), \(y=1-s\)): the positive-definiteness restriction of Theorem 20.6.10 is lifted. The general two-exponent region is obtained in Theorem 20.6.12.

Proof

For the interior exponents \(s\in (0,1)\), replace each input \(A_i\) by the positive-definite matrix \(A_i+\varepsilon \mathbb {1}\), and likewise for the \(B_i\). Theorem 20.6.10 applies to these positive-definite matrices for every \(\varepsilon {\gt}0\), and the average of the regularized inputs is the regularization of the average. Letting \(\varepsilon \to 0^+\), each regularized power \((A_i+\varepsilon \mathbb {1})^s\) converges to \(A_i^s\), since \(x\mapsto x^s\) is continuous on the non-negative reals for \(0{\lt}s{\lt}1\), including at the origin; the trace functional is continuous, so the regularized inequality passes to the limit. The endpoints \(s=0\) and \(s=1\) are linear in the convex-combination argument and hold with equality.

Theorem 20.6.12 Lieb concavity, sub-boundary exponents

Let \(x,y\ge 0\) with \(x+y\le 1\). For any matrix \(K\) and positive-semidefinite matrices \(A_1,A_2,B_1,B_2\), write \(F_{x,y}(A,B):=\Re \operatorname{tr}(K^\dagger A^x K B^y)\). This map is jointly concave. For \(t\in [0,1]\), set \(A_t:=tA_1+(1-t)A_2\) and \(B_t:=tB_1+(1-t)B_2\). Then

\begin{align} tF_{x,y}(A_1,B_1)+(1-t)F_{x,y}(A_2,B_2) & \le F_{x,y}(A_t,B_t). \notag \end{align}

This is Wolf’s Theorem 5.15 in the finite-dimensional matrix setting.

Proof

Put \(r=x+y\). If \(r=0\), both exponents vanish and the functional is constant. If \(0{\lt}r\le 1\), set \(s=x/r\). Then \(x=rs\) and \(y=r(1-s)\), with \(s\in [0,1]\). The functional becomes the boundary Lieb functional applied to the powered inputs \(A^r\) and \(B^r\). Operator concavity of \(C\mapsto C^r\) gives

\begin{align} tA_1^r+(1-t)A_2^r & \le (tA_1+(1-t)A_2)^r, \notag \\ tB_1^r+(1-t)B_2^r & \le (tB_1+(1-t)B_2)^r. \notag \end{align}

The boundary positive-semidefinite theorem gives concavity at the left-hand powered inputs, and monotonicity of \((P,Q)\mapsto \Re \operatorname{tr}(K^\dagger P^s K Q^{1-s})\) on positive-semidefinite arguments transports the result to the powered convex combinations.

Theorem 20.6.13 Map form of Lieb concavity
#

For \(s\in [0,1]\), any matrix \(K\), and positive-definite matrices \(A_1,A_2,B_1,B_2\), the function \((A,B)\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is jointly concave, satisfying (??) of Theorem 20.6.10.

Proof

Follows directly from Theorem 20.6.10.

Corollary 20.6.14 Lieb concavity, \(K=\mathbb {1}\) case
#

For \(s\in [0,1]\), \(t\in [0,1]\), and positive-definite matrices \(A_1,A_2,B_1,B_2\), the map \((A,B)\mapsto \Re \operatorname{tr}(A^s B^{1-s})\) is jointly concave:

\begin{align} t\, \Re \operatorname{tr}(A_1^s B_1^{1-s}) +(1-t)\Re \operatorname{tr}(A_2^s B_2^{1-s}) & \le \Re \operatorname{tr}\! \bigl( (tA_1+(1-t)A_2)^s (tB_1+(1-t)B_2)^{1-s} \bigr). \notag \end{align}

Obtained from Theorem 20.6.13 by taking \(K=\mathbb {1}\).

Proof

Specialize Theorem 20.6.13 at \(K=\mathbb {1}\) and simplify \(\mathbb {1}^\dagger =\mathbb {1}\) together with \(\mathbb {1}\cdot X=X\cdot \mathbb {1}=X\).

Corollary 20.6.15 Concavity of Lieb’s map in the first argument
#

For \(s\in [0,1]\), \(t\in [0,1]\), matrix \(K\), and positive-definite matrices \(A_1,A_2,B\), the map \(A\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is concave:

\begin{align} t\, \Re \operatorname{tr}(K^\dagger A_1^s K B^{1-s}) +(1-t)\Re \operatorname{tr}(K^\dagger A_2^s K B^{1-s}) & \le \Re \operatorname{tr}\! \bigl( K^\dagger (tA_1+(1-t)A_2)^s K B^{1-s} \bigr). \notag \end{align}
Proof

Specialize Theorem 20.6.13 at \(B_1=B_2=B\), so that the convex combination collapses: \(tB+(1-t)B=B\).

Corollary 20.6.16 Concavity of Lieb’s map in the second argument
#

For \(s\in [0,1]\), \(t\in [0,1]\), matrix \(K\), and positive-definite matrices \(A,B_1,B_2\), the map \(B\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is concave:

\begin{align} t\, \Re \operatorname{tr}(K^\dagger A^s K B_1^{1-s}) +(1-t)\Re \operatorname{tr}(K^\dagger A^s K B_2^{1-s}) & \le \Re \operatorname{tr}\! \bigl( K^\dagger A^s K(tB_1+(1-t)B_2)^{1-s} \bigr). \notag \end{align}
Proof

Specialize Theorem 20.6.13 at \(A_1=A_2=A\), so that the convex combination collapses: \(tA+(1-t)A=A\).

20.7 Resolvent monotonicity toward the Lieb concavity theorem

The integral-representation route toward the Lieb concavity theorem rests on the antitonicity of the resolvent of the commuting left- and right-multiplication superoperators. Its matrix-order foundation is the inverse-antitonicity of the Loewner order.

Lemma 20.7.1 Loewner inverse-antitonicity
#

For positive-definite matrices \(A\) and \(B\) with \(A\le B\) in the Loewner order, the inverses satisfy \(B^{-1}\le A^{-1}\).

Proof

Compare the two Schur complements of the block matrix \(\begin{psmallmatrix} \end{psmallmatrix}A^{-1}& \mathbb {1}\\ \mathbb {1}& B\end{psmallmatrix}\). Complementing its \((1,1)\) block gives \(B-(A^{-1})^{-1}=B-A\ge 0\), so the block matrix is positive semidefinite; complementing the \((2,2)\) block of the same matrix gives \(A^{-1}-B^{-1}\), which is therefore positive semidefinite as well.

Lemma 20.7.2 Joint Loewner-antitonicity of the commuting-multiplication resolvent
#

Fix \(t{\gt}0\) and positive-definite matrices with \(A_1\le A_2\) and \(B_1\le B_2\). The resolvent of the Kronecker model \(A\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B^\top )\) of the commuting left- and right-multiplication superoperators is antitone:

\begin{align} (A_2\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B_2^\top ))^{-1} & \le (A_1\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B_1^\top ))^{-1}. \notag \end{align}
Proof

Each operator \(A\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B^\top )\) is positive definite, being the sum of the positive-definite \(A\otimes \mathbb {1}\) and the positive-semidefinite \(t\, (\mathbb {1}\otimes B^\top )\). From \(A_1\le A_2\) and \(B_1\le B_2\) one gets \(A_1\otimes \mathbb {1}\le A_2\otimes \mathbb {1}\) and \(\mathbb {1}\otimes B_1^\top \le \mathbb {1}\otimes B_2^\top \); since \(t{\gt}0\), adding these yields

\begin{align} A_1\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B_1^\top ) & \le A_2\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B_2^\top ). \notag \end{align}

Inverse-antitonicity (Lemma 20.7.1) reverses this inequality on passing to the resolvents.