20 Operator Convexity and Jensen Inequalities
This chapter collects the operator convexity/concavity results, trace inequalities, the operator Jensen inequality for positive maps, and the Lieb concavity theorem. These are standard results in matrix analysis ( [ Bha97 , Chapter V ] ; [ Wol12 , Theorems 5.12 and 5.13 ] ; Hansen–Pedersen’s Jensen inequality [ HP82 ] ; and Lieb’s concavity theorem [ Lie73 ] ). The operator Jensen inequalities for real powers \(x\mapsto x^p\) on the Loewner order are obtained from the Loewner integral representation of the power function. Lieb’s concavity theorem is derived from the operator integral representation of the fractional product together with the vectorization isometry, which turns the trace functional into a quadratic form of a positive-semidefinite matrix.
20.1 Schwarz inequalities for positive maps
The following consequences of positivity and the Kadison–Schwarz inequality concern normal, subnormal, and commuting-dominant operators. They provide matrix-order tools used independently of the Fundamental Theorem.
20.1.1 Schwarz inequality for normal operators
Let \(E^*(X)=\sum _i K_i^\dagger X K_i\) be the adjoint Kraus map of a trace-preserving Kraus family, and let \(A\in M_{D}(\mathbb {C})\) be normal. Then the Schwarz gap
is positive semidefinite. Equivalently,
This is the completely positive case of [ Wol12 , Proposition 5.1 ] .
Positivity of the gap in (??) is exactly Theorem 5.1.6. The Loewner inequality (??) is its order-theoretic reformulation.
Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be a positive linear map with \(T(\mathbb {1})\le \mathbb {1}\). If \(A\in M_{D}(\mathbb {C})\) is normal, then \(T(A^\dagger A)-T(A^\dagger )T(A)\ge 0\). This is [ Wol12 , Proposition 5.1 ] .
Since \(A\) is normal, diagonalize \(A=UDU^\dagger \) with \(D\) diagonal and \(U\) unitary. The spectral projectors of \(A\) generate a commutative \(*\)-subalgebra. The restriction of \(T\) to this subalgebra is completely positive, so the Kadison–Schwarz inequality for completely positive subunital maps gives \(T(A^\dagger )T(A)\le T(A^\dagger A)\).
The proof follows the positivity-on-abelian argument of [ Wol12 , Proposition 1.6 ] : a positive map restricted to the commutative \(*\)-subalgebra generated by the spectral projectors of \(A\) satisfies the Schwarz inequality directly, without invoking the Arveson extension theorem.
20.1.2 Schwarz inequality for subnormal and commuting-dominant operators
Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be a positive linear map with \(T(\mathbb {1})\le \mathbb {1}\). If \(A\in M_{D}(\mathbb {C})\) is subnormal (i.e. there exists a normal operator on a larger space whose compression to \(\mathbb {C}^D\) is \(A\)), then \(T(A^\dagger A)-T(A^\dagger )T(A)\ge 0\). This is [ Wol12 , Theorem 5.5 ] .
Embed \(A\) as the northwest corner of a normal operator \(N\) on \(\mathbb {C}^D\oplus \mathbb {C}^E\). Extend \(T\) to a positive subunital map \(\widetilde{T}\) on the larger space, acting as \(T\) on the northwest block and as zero elsewhere. Apply Theorem 20.1.1.2 to \(\widetilde{T}\) and \(N\), then compress the resulting inequality back to \(\mathbb {C}^D\).
Let \(T:M_{D}(\mathbb {C})\to M_{D}(\mathbb {C})\) be a positive linear map with \(T(\mathbb {1})\le \mathbb {1}\). If \(A\in M_{D}(\mathbb {C})\) admits a positive semidefinite dominant \(B\ge 0\) with \(A^\dagger A\le B\) and \(B\) commuting with \(A\) (\(BA=AB\)), then
This is [ Wol12 , Theorem 5.6 ] .
First take \(B{\gt}0\). Since \(B\) commutes with \(A\) and \(A^\dagger A\le B\), also \(AA^\dagger \le B\), so \(D_R=(B-A^\dagger A)^{1/2}\) and \(D_L=(B-AA^\dagger )^{1/2}\) are positive semidefinite. The block matrix
is normal. As in the subnormal case, extend \(T\) to a positive subunital map on \(M_{2D}(\mathbb {C})\), apply the normal-operator Schwarz inequality (Theorem 20.1.1.2) to that extension and \(N\), and compress back to \(\mathbb {C}^D\). Since \(N\) is normal, both \(T(A^\dagger )T(A)\le T(B)\) and \(T(A)T(A^\dagger )\le T(B)\) follow. For general \(B\ge 0\), apply this to \(B_\varepsilon =B+\varepsilon \mathbb {1}{\gt}0\) and let \(\varepsilon \to 0\).
Let \(E^*\) be the adjoint of a trace-preserving Kraus map. If \(A\in M_{D}(\mathbb {C})\) admits a positive semidefinite dominant \(B\) commuting with \(A\) and satisfying \(A^\dagger A\le B\), then
Apply Theorem 20.1.2.2 to the adjoint Kraus map, which is unital under the trace-preservation hypothesis.
20.2 Diagonal Jensen inequality
Let \(f:\mathbb {R}\to \mathbb {R}\) be convex on \([0,\infty )\), let \(A\) be a positive semidefinite matrix on a finite-dimensional space indexed by \(n\), and let \(v\in \mathbb {C}^n\) satisfy \(\langle v,v\rangle =1\). Then
where \(f(A)\) is defined by the Hermitian continuous functional calculus.
Write \(A=U\operatorname{diag}(\mu )U^\dagger \) by the spectral theorem and set \(w=U^\dagger v\). Since \(U\) is unitary, \(p_i:=\lvert w_i\rvert ^2\) satisfies \(\sum _i p_i=\lVert v\rVert ^2=1\), so \((p_i)\) is a probability distribution on \(n\). The eigenvalues \(\mu _i\) lie in \([0,\infty )\) because \(A\) is positive semidefinite. A direct computation yields
so the scalar Jensen inequality applied to the weights \((p_i)\) and points \((\mu _i)\) gives the conclusion. See [ Bha97 , Chapter V ] .
20.3 Trace concavity and convexity of matrix powers
For \(p\in [0,1]\), PSD matrices \(A_1,A_2\), and \(t\in [0,1]\),
This follows from operator concavity of \(x\mapsto x^p\) composed with trace monotonicity on the Loewner order.
Let \(A=tA_1+(1-t)A_2\), and let \(\{ \psi _j\} \) be an orthonormal eigenbasis of \(A\) with eigenvalues \(\mu _j\ge 0\). A direct computation gives
Scalar concavity of \(x\mapsto x^p\) on \([0,\infty )\) yields \(\mu _j^p\ge t\, a_j^p+(1-t)b_j^p\). Applying Lemma 20.2.1 to the convex function \(x\mapsto -x^p\) and multiplying by \(-1\) gives
Summing over \(j\) and using \(\sum _j\operatorname{Re}\langle \psi _j,M\psi _j\rangle =\operatorname{Re}\operatorname{tr}M\) yields the conclusion. See [ Bha97 , Chapter V ] .
For \(p\in [1,2]\), PSD matrices \(A_1,A_2\), and \(t\in [0,1]\),
This follows from operator convexity of \(x\mapsto x^p\) composed with trace monotonicity on the Loewner order.
Use the eigenbasis reduction from Theorem 20.3.1, with scalar convexity rather than concavity of \(x\mapsto x^p\) on \([0,\infty )\) for \(p\ge 1\) and the convex form of the diagonal Jensen inequality.
20.4 Operator convexity of real powers
For elements \(a,b\) of a unital C\({}^\ast \)-algebra with \(0\le a,b\), an exponent \(p\in [1,2]\), and \(t\in [0,1]\), the map \(a\mapsto a^p\) is convex for the Loewner order on the positive cone:
This is the convex counterpart of the operator concavity of \(x\mapsto x^p\) for exponents \(p\in [0,1]\).
For \(p\in (1,2)\) one uses the integral representation \(a^p=\int _0^\infty f_{p,s}(a)\, d\mu (s)\) of Carlen [ Car10 , Lemma 2.8 ] , whose integrand \(f_{p,s}(x)=s^{p-1}(s^{-1}x+s(s+x)^{-1}-1)\) is operator convex: it decomposes into a linear term, a non-negative multiple of the operator convex resolvent \(x\mapsto (s\cdot 1+x)^{-1}\), and a constant. Since \(f_{p,s}(ta+(1-t)b)\le t\, f_{p,s}(a)+(1-t)f_{p,s}(b)\) for every \(s{\gt}0\), integrating this pointwise Loewner inequality against \(\mu \) gives
The endpoint \(p=1\) is the identity, and \(p=2\) is \(a\mapsto a^2\), operator convex because \(t\, a^2+(1-t)b^2-(ta+(1-t)b)^2=t(1-t)(a-b)^2\ge 0\). See [ Bha97 , Chapter V ] .
20.5 Operator Jensen inequality for positive maps
For a finite family of matrices \(\{ C_i\} _{i\in \iota }\) and a defect block \(S\), one forms an isometric dilation whose rows are the adjoints of the matrices \(C_i\) together with the adjoint of \(S\). One also forms the scalar block-diagonal matrix whose \(\iota \)-blocks carry prescribed weights and whose defect block carries a single scalar.
Let \(Y\) be positive definite and let \(W\) be an isometry. Then
Apply the Schur-complement criterion to the block matrix \(\begin{psmallmatrix} \end{psmallmatrix}Y^{-1}& W\\ W^\dagger & W^\dagger YW\end{psmallmatrix}\). The lower-right Schur complement is zero, hence positive semidefinite, and compressing the resulting upper-left Schur complement by \(W\) gives (??).
If the defect block satisfies
then the dilation is an isometry, and compressing the weighted scalar block-diagonal matrix gives \(\sum _{i\in \iota }w_iC_iC_i^\dagger +tSS^\dagger \). The defect relation can also be written as \(\sum _{i\in \iota }C_iC_i^\dagger +SS^\dagger =\mathbb {1}\). Strict positivity of all weights together with \(t{\gt}0\) implies positive definiteness of the scalar block-diagonal matrix. If all weights are nonzero and the defect scalar is nonzero, the inverse diagonal has reciprocal weights, and compressing this inverse gives \(\sum _{i\in \iota }w_i^{-1}C_iC_i^\dagger +t^{-1}SS^\dagger \).
Expand the block matrix multiplication entrywise. The defect identity (??) gives the isometry relation after summing the \(\iota \)-blocks and the defect block, and the weighted compression identity follows from the same expansion with the scalar diagonal inserted. Positive definiteness uses strict positivity of every diagonal entry, while the inverse formula uses the nonvanishing of every diagonal entry.
Let \(w_i\ge 0\) and \(t{\gt}0\). If the defect block satisfies \(SS^\dagger =\mathbb {1}-\sum _{i\in \iota }C_iC_i^\dagger \), then
Apply the compression inequality (??) to the positive definite scalar block-diagonal matrix with entries \(w_i+t\) and defect entry \(t\). The compression identities rewrite the compressed matrix as \(\sum _iw_iC_iC_i^\dagger +t\mathbb {1}\) and its inverse compression as the upper bound in (??).
Let \(A\ge 0\) be a finite matrix. There are positive semidefinite matrices \(P_i\) and non-negative real numbers \(\lambda _i\) such that
Diagonalize the Hermitian matrix \(A\) by a unitary matrix. The matrices \(P_i\) are the conjugates of the rank-one diagonal matrix units, their sum is the identity, and positive semidefiniteness of \(A\) gives \(\lambda _i\ge 0\).
Let \(T\) be a positive subunital map, let \(A\ge 0\), and let \(t{\gt}0\). Then
Use the spectral resolution \(A=\sum _i\lambda _iP_i\). The matrices \(T(P_i)\) are positive semidefinite and sum to \(T(\mathbb {1})\le \mathbb {1}\). Write each \(T(P_i)\) as \(C_iC_i^\dagger \) using the positive square root, and write the defect \(\mathbb {1}-\sum _iT(P_i)\) as \(SS^\dagger \). The finite-POVM resolvent inequality gives the desired estimate after substituting the spectral formula for \((A+t\mathbb {1})^{-1}\).
Let \(T\) be a positive subunital map, let \(A\ge 0\), let \(p\in (0,1)\), and let \(t{\gt}0\). Then
Rewrite the continuous functional calculus in resolvent form as \(f_{p,t}(X)=t^{p-1}\mathbb {1}-t^p(t\mathbb {1}+X)^{-1}\). After applying the linear map \(T\), the difference \(f_{p,t}(T(A))-T(f_{p,t}(A))\) is exactly \(t^p\) times the positive matrix supplied by the positive-map resolvent inequality.
Let \(T\) be a positive subunital map, let \(A\ge 0\), let \(p\in (1,2)\), and let \(t{\gt}0\). Then
The inequality is reversed relative to the concave integrand, since \(g_{p,t}\) is operator convex.
Rewrite the continuous functional calculus in resolvent form as \(g_{p,t}(X)=t^{p-2}X+t^p(t\mathbb {1}+X)^{-1}-t^{p-1}\mathbb {1}\). Under the linear map \(T\), the leading \(t^{p-2}\) terms cancel and the constant \(t^{p-1}\) terms combine, so the difference \(T(g_{p,t}(A))-g_{p,t}(T(A))\) is exactly \(t^p\) times the positive matrix supplied by the positive-map resolvent inequality, using \(t^pt^{-1}=t^{p-1}\).
Let \(f\) be a function from a measure space to a finite matrix algebra. If \(f(x)\ge 0\) for almost every \(x\), then \(\int f(x)\, d\mu (x)\ge 0\).
The Loewner order has closed positive cone in finite dimension. Hence the ordered Bochner integral preserves almost-everywhere nonnegativity. The corresponding almost-everywhere monotonicity statement is the general ordered Bochner integral theorem.
For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [0,1]\), one has \(T(A^p)\le (T(A))^p\). This is [ Wol12 , Theorem 5.13 ] for concave \(f(x)=x^p\).
For \(p=0\) the assertion is the subunitality condition \(T(\mathbb {1})\le \mathbb {1}\). For \(p=1\) it is an equality. For \(0{\lt}p{\lt}1\), use Carlen’s integral representation of \(A^p\) by the functions \(f_{p,t}(x)=t^p(t^{-1}-(t+x)^{-1})\). The pointwise inequality \(T(f_{p,t}(A))\le f_{p,t}(T(A))\) is Lemma 20.5.7. Integrating this inequality with respect to the representing measure gives the result by Lemma 20.5.9 and linearity of the Bochner integral under \(T\).
For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [0,1]\), the inequality \(T(A^p)\le (T(A))^p\) holds.
Direct application of Theorem 20.5.10.
For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [1,2]\), one has \((T(A))^p\le T(A^p)\). This is [ Wol12 , Theorem 5.11 ] for the subunital operator-convex case \(f(x)=x^p\).
For \(p=1\) the assertion is an equality. For \(1{\lt}p{\lt}2\), use Carlen’s integral representation of \(A^p\) by the convex integrands \(g_{p,t}(x)=t^{p-2}x+t^p(t+x)^{-1}-t^{p-1}\). The pointwise inequality \(g_{p,t}(T(A))\le T(g_{p,t}(A))\) is Lemma 20.5.8. Integrating this inequality with respect to the representing measure and using Lemma 20.5.9 together with linearity of the Bochner integral under \(T\) gives the inequality for \(1{\lt}p{\lt}2\). The endpoint \(p=2\) follows by taking the limit \(q\to 2^-\) in \((T(A))^q\le T(A^q)\): since \(M^q\to M^2\) in the operator norm as \(q\to 2^-\) for any positive semidefinite \(M\), and the Loewner order is closed, this gives \((T(A))^2\le T(A^2)\).
For a positive subunital map \(T\), a positive semidefinite matrix \(A\), and \(p\in [1,2]\), the inequality \((T(A))^p\le T(A^p)\) holds.
Direct application of Theorem 20.5.12.
For a positive unital map \(T\) and positive-definite \(A\), one has \(T(\log A)\le \log (T(A))\). This is [ Wol12 , Theorem 5.13 ] for \(f=\log \). Requires unitality (\(T(\mathbb {1})=\mathbb {1}\)), not merely subunitality.
Apply Theorem 20.5.10 to \(0{\lt}p{\lt}1\) and rewrite it as an inequality for \((A^p-\mathbb {1})/p\). The continuous functional calculus gives \((A^p-\mathbb {1})/p\to \log A\) as \(p\to 0^+\), and the same limit holds for \(T(A)\). Since the Loewner order graph is closed in finite dimension, the limiting inequality is \(T(\log A)\le \log (T(A))\).
For a positive unital map \(T\) and a positive-definite matrix \(A\), the inequality \(T(\log A)\le \log (T(A))\) holds.
Direct application of Theorem 20.5.14.
20.6 Lieb concavity theorem
The Lieb concavity theorem rests on a scalar integral identity that expresses the geometric interpolation \(a^s b^{1-s}\) as a positive average of the elementary fractions \(ab/(a+tb)\). This is the analytic core of the integral representation of \(A^s B^{1-s}\).
For \(s\in (0,1)\),
This is the value \(B(s,1-s)\) of the Beta function, which equals \(\Gamma (s)\Gamma (1-s)=\pi /\sin (\pi s)\) by Euler’s reflection formula. The substitution \(x=u/(1+u)\) carries the half-line onto the unit interval and turns the integrand into the Beta integrand \(x^{s-1}(1-x)^{-s}\).
For \(a,b{\gt}0\) and \(s\in (0,1)\), the function \(t\mapsto t^{s-1} \dfrac {ab}{a+tb}\) is integrable on \((0,\infty )\).
The integrand is \(O(t^{s-1})\) as \(t\to 0^+\) and \(O(t^{s-2})\) as \(t\to \infty \), so it has integrable comparison functions on each tail.
For \(a,b{\gt}0\) and \(s\in (0,1)\),
The scaling substitution \(t=(a/b)\, u\) turns the integrand into \(a^s b^{1-s}\, u^{s-1}/(1+u)\), whose half-line integral is the reflection integral \(\pi /\sin (\pi s)\) of Lemma 20.6.1.
For \(a,b{\gt}0\) and \(s\in (0,1)\),
By Lemma 20.6.3, the integral equals \(a^s b^{1-s} \pi /\sin (\pi s)\), so
For positive diagonal matrices \(M\) and \(N\) and \(s\in (0,1)\),
Both sides are diagonal, so the identity holds entry by entry; on each diagonal entry it is the scalar Lieb integral identity (Lemma 20.6.4).
For positive-definite matrices \(A\), \(B\) and \(s\in (0,1)\),
where \(A\otimes \mathbb {1}\) and \(\mathbb {1}\otimes B^\top \) are regarded as operators on the Kronecker model space \(\mathbb {C}^{D\times D}\).
The pair \((A\otimes \mathbb {1},\mathbb {1}\otimes B^\top )\) is simultaneously diagonalized by \(V=U_A\otimes U_{B^\top }\) (eigenbases of \(A\) and \(B^\top \)). In that basis the pair becomes a commuting pair of positive diagonal matrices, and the identity follows from the diagonal case (Lemma 20.6.5).
For positive-definite matrices \(X_1,X_2,Y_1,Y_2\) and \(\theta \in [0,1]\), the parallel sum \((X,Y)\mapsto X(X+Y)^{-1}Y\) is jointly concave in the Loewner order:
where \(X_\theta =\theta X_1+(1-\theta )X_2\) and \(Y_\theta =\theta Y_1+(1-\theta )Y_2\).
Write \(Z=X(X+Y)^{-1}Y\). The identity \(X(X+Y)^{-1}(X+Y)=X\) gives \((X-Z)-X(X+Y)^{-1}X=0\), so the block matrix \(\begin{psmallmatrix} \end{psmallmatrix}X-Z & X\\ X & X+Y\end{psmallmatrix}\) is positive semidefinite (its Schur complement against \(X+Y\) vanishes). Setting \(Z_i=X_i(X_i+Y_i)^{-1}Y_i\) and \(Z_\theta =\theta Z_1+(1-\theta )Z_2\), the convex combination of the two block matrices is \(\begin{psmallmatrix} \end{psmallmatrix}X_\theta -Z_\theta & X_\theta \\ X_\theta & X_\theta +Y_\theta \end{psmallmatrix}\), which is positive semidefinite, and its Schur complement against \(X_\theta +Y_\theta \) yields \(X_\theta (X_\theta +Y_\theta )^{-1}Y_\theta -Z_\theta \ge 0\), which is the claimed inequality (Ando, 1979).
For \(t{\gt}0\), positive-definite \(A_1,A_2,B_1,B_2\), and \(\theta \in [0,1]\), with \(\hat{A}=A\otimes \mathbb {1}\) and \(\hat{B}=\mathbb {1}\otimes B^\top \), setting \(A_\theta =\theta A_1+(1-\theta )A_2\) and \(B_\theta =\theta B_1+(1-\theta )B_2\):
The integrand equals \(t^{-1} \hat A(\hat A+t\hat B)^{-1}(t\hat B)\), a rescaled parallel sum of \(\hat A\) and \(t\hat B\), so concavity follows from Lemma 20.6.7 and the linearity of \(A\mapsto \hat A\), \(B\mapsto \hat B\).
For positive-definite matrices \(A_1,A_2,B_1,B_2\), \(s\in (0,1)\), and \(\theta \in [0,1]\), with \(\hat{A}=A\otimes \mathbb {1}\) and \(\hat{B}=\mathbb {1}\otimes B^\top \), setting \(A_\theta =\theta A_1+(1-\theta )A_2\) and \(B_\theta =\theta B_1+(1-\theta )B_2\), the fractional product \(\hat A^s\hat B^{1-s}\) is jointly concave in the Loewner order:
By the integral representation (??), each fractional product equals the positive multiple \(\frac{\sin (\pi s)}{\pi }\) of the integral of the resolvent integrand against the non-negative weight \(t^{s-1}\) on \((0,\infty )\). For each \(t{\gt}0\), the integrand is jointly concave by Lemma 20.6.8, so the difference of the averaged point and the convex combination is, almost everywhere, the non-negative scalar \(t^{s-1}\) times a positive-semidefinite matrix. Integrating a positive-semidefinite integrand yields a positive-semidefinite matrix, and scaling by the positive constant \(\frac{\sin (\pi s)}{\pi }\) preserves the Loewner order, which gives the inequality.
For \(s\in [0,1]\), any matrix \(K\), and positive-definite matrices \(A_1,A_2,B_1,B_2\), write \(F_s(A,B):=\Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\). This map is jointly concave. For \(t\in [0,1]\), set \(A_t:=tA_1+(1-t)A_2\) and \(B_t:=tB_1+(1-t)B_2\). Then
See [ Lie73 ; And79 ] . This statement is the positive-definite, boundary-exponent (\(x+y=1\)) case of the Ando–Lieb theorem [ Wol12 , Theorem 5.15 ] , which holds for positive semidefinite \(A,B\) and all exponents \(x,y\ge 0\) with \(x+y\le 1\).
For \(s\in (0,1)\), the concavity of the fractional product of the commuting left and right multiplication operators (Lemma 20.6.9) gives a Loewner-order inequality between positive-semidefinite matrices on the model space \(\mathbb {C}^D\otimes \mathbb {C}^D\). Writing that product as \(A^s\otimes (B^\top )^{1-s}\) and pairing it with the vectorization of \(K^\top \),
expresses the trace functional as the quadratic form of the Kronecker product at that vector. Applying this positive quadratic form to the Loewner inequality transfers the concavity to the real part of the trace. The endpoints \(s=0\) and \(s=1\) reduce the varying side to a linear function of the convex-combination argument, hence hold with equality.
For \(s\in [0,1]\), any matrix \(K\), and positive-semidefinite matrices \(A_1,A_2,B_1,B_2\), the map \((A,B)\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is jointly concave, satisfying (??) of Theorem 20.6.10. This is the full Ando–Lieb theorem [ Wol12 , Theorem 5.15 ] on the boundary line \(x+y=1\) (with \(x=s\), \(y=1-s\)): the positive-definiteness restriction of Theorem 20.6.10 is lifted. The general two-exponent region is obtained in Theorem 20.6.12.
For the interior exponents \(s\in (0,1)\), replace each input \(A_i\) by the positive-definite matrix \(A_i+\varepsilon \mathbb {1}\), and likewise for the \(B_i\). Theorem 20.6.10 applies to these positive-definite matrices for every \(\varepsilon {\gt}0\), and the average of the regularized inputs is the regularization of the average. Letting \(\varepsilon \to 0^+\), each regularized power \((A_i+\varepsilon \mathbb {1})^s\) converges to \(A_i^s\), since \(x\mapsto x^s\) is continuous on the non-negative reals for \(0{\lt}s{\lt}1\), including at the origin; the trace functional is continuous, so the regularized inequality passes to the limit. The endpoints \(s=0\) and \(s=1\) are linear in the convex-combination argument and hold with equality.
Let \(x,y\ge 0\) with \(x+y\le 1\). For any matrix \(K\) and positive-semidefinite matrices \(A_1,A_2,B_1,B_2\), write \(F_{x,y}(A,B):=\Re \operatorname{tr}(K^\dagger A^x K B^y)\). This map is jointly concave. For \(t\in [0,1]\), set \(A_t:=tA_1+(1-t)A_2\) and \(B_t:=tB_1+(1-t)B_2\). Then
This is Wolf’s Theorem 5.15 in the finite-dimensional matrix setting.
Put \(r=x+y\). If \(r=0\), both exponents vanish and the functional is constant. If \(0{\lt}r\le 1\), set \(s=x/r\). Then \(x=rs\) and \(y=r(1-s)\), with \(s\in [0,1]\). The functional becomes the boundary Lieb functional applied to the powered inputs \(A^r\) and \(B^r\). Operator concavity of \(C\mapsto C^r\) gives
The boundary positive-semidefinite theorem gives concavity at the left-hand powered inputs, and monotonicity of \((P,Q)\mapsto \Re \operatorname{tr}(K^\dagger P^s K Q^{1-s})\) on positive-semidefinite arguments transports the result to the powered convex combinations.
For \(s\in [0,1]\), any matrix \(K\), and positive-definite matrices \(A_1,A_2,B_1,B_2\), the function \((A,B)\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is jointly concave, satisfying (??) of Theorem 20.6.10.
Follows directly from Theorem 20.6.10.
For \(s\in [0,1]\), \(t\in [0,1]\), and positive-definite matrices \(A_1,A_2,B_1,B_2\), the map \((A,B)\mapsto \Re \operatorname{tr}(A^s B^{1-s})\) is jointly concave:
Obtained from Theorem 20.6.13 by taking \(K=\mathbb {1}\).
Specialize Theorem 20.6.13 at \(K=\mathbb {1}\) and simplify \(\mathbb {1}^\dagger =\mathbb {1}\) together with \(\mathbb {1}\cdot X=X\cdot \mathbb {1}=X\).
For \(s\in [0,1]\), \(t\in [0,1]\), matrix \(K\), and positive-definite matrices \(A_1,A_2,B\), the map \(A\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is concave:
Specialize Theorem 20.6.13 at \(B_1=B_2=B\), so that the convex combination collapses: \(tB+(1-t)B=B\).
For \(s\in [0,1]\), \(t\in [0,1]\), matrix \(K\), and positive-definite matrices \(A,B_1,B_2\), the map \(B\mapsto \Re \operatorname{tr}(K^\dagger A^s K B^{1-s})\) is concave:
Specialize Theorem 20.6.13 at \(A_1=A_2=A\), so that the convex combination collapses: \(tA+(1-t)A=A\).
20.7 Resolvent monotonicity toward the Lieb concavity theorem
The integral-representation route toward the Lieb concavity theorem rests on the antitonicity of the resolvent of the commuting left- and right-multiplication superoperators. Its matrix-order foundation is the inverse-antitonicity of the Loewner order.
For positive-definite matrices \(A\) and \(B\) with \(A\le B\) in the Loewner order, the inverses satisfy \(B^{-1}\le A^{-1}\).
Compare the two Schur complements of the block matrix \(\begin{psmallmatrix} \end{psmallmatrix}A^{-1}& \mathbb {1}\\ \mathbb {1}& B\end{psmallmatrix}\). Complementing its \((1,1)\) block gives \(B-(A^{-1})^{-1}=B-A\ge 0\), so the block matrix is positive semidefinite; complementing the \((2,2)\) block of the same matrix gives \(A^{-1}-B^{-1}\), which is therefore positive semidefinite as well.
Fix \(t{\gt}0\) and positive-definite matrices with \(A_1\le A_2\) and \(B_1\le B_2\). The resolvent of the Kronecker model \(A\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B^\top )\) of the commuting left- and right-multiplication superoperators is antitone:
Each operator \(A\otimes \mathbb {1}+t\, (\mathbb {1}\otimes B^\top )\) is positive definite, being the sum of the positive-definite \(A\otimes \mathbb {1}\) and the positive-semidefinite \(t\, (\mathbb {1}\otimes B^\top )\). From \(A_1\le A_2\) and \(B_1\le B_2\) one gets \(A_1\otimes \mathbb {1}\le A_2\otimes \mathbb {1}\) and \(\mathbb {1}\otimes B_1^\top \le \mathbb {1}\otimes B_2^\top \); since \(t{\gt}0\), adding these yields
Inverse-antitonicity (Lemma 20.7.1) reverses this inequality on passing to the resolvents.