From Calculus to Kelly量化數理教科書 Vol I · LM7 Multivariable Calculus and Lagrange Multipliers
Learning Module
7

Multivariable Calculus and Lagrange Multipliers

Gradients, Hessians and multipliers: how calculus finds the best mix of two assets, and what that answer is worth out of sample.

Volume I · Calculus for FinanceBuilds on LM2 · LM3 · LM4Leads to LM10 · LM32 · LM33 · LM34 · LM43Time ≈ 3 h + labPDF A4 print edition
Learning Outcomes
MasteryAfter this module you should be able to:
1compute partial derivatives and total differentials of functions of two or three variables, and interpret them as sensitivities with the other inputs held fixed
2construct the gradient and Hessian of a function of several variables, and approximate the function with its second-order multivariate Taylor expansion
3locate and classify the critical points of a function of two variables with the first-order conditions and the 2×2 definiteness test
4solve equality-constrained problems with Lagrange multipliers, and interpret the multiplier as a shadow price
5derive the two-asset minimum-variance weight and explain how correlation determines the diversification benefit
6derive the two-asset tangency weights and Kelly fractions, and relate them to
7apply the Karush–Kuhn–Tucker conditions to long-only, policy and leverage constraints, and identify which constraints bind
8implement numerical gradients, Newton's method and constrained optimization in Python, and distinguish in-sample optimal weights from out-of-sample results

1Introduction

Every portfolio decision turns more than one dial: how much 0050, how much bond, how much leverage. Single-variable calculus (LM2, LM4) turns one dial at a time; this module turns them together, adds constraints and prices the constraints themselves. Applied to two assets, it yields the minimum-variance portfolio, the tangency portfolio and the Kelly fractions, and it shows why an answer fitted to the past is not the answer you will get.

Motivating CaseHow much bond minimizes risk?

A TWD investor holds Yuanta Taiwan 50 (0050) and considers adding Yuanta U.S. Treasury 20+ Year Bond ETF (00679B). The fund tracks the ICE US Treasury 20+ Year Bond Index and is quoted in New Taiwan dollars. Its holdings report lists US Treasuries and US-dollar cash but no currency forwards, so its returns combine US long-bond returns with USD/TWD moves.

From the close of 17 January 2017, 00679B's first trading day, to 8 October 2026, 0050 returned +767.3% with dividends reinvested, at 20.1% annualized volatility. 00679B returned −17.3% at 13.5%. Their daily returns were mildly negatively correlated, −0.11 (Exhibit 1).

Suppose the investor had sized the mix at the end of 2021 with the five years then available. Over 2017–2021 the correlation was −0.22, and the calculus of Section 6 says that 42.3% in 0050 and 57.7% in 00679B minimized volatility, at a predicted 9.22% a year.

Then came 2022: 0050 lost 21.3% and 00679B lost 22.1%. Over 2022–2026 the correlation drifted to −0.02, and the same mix realized 12.44%. Was the calculus wrong, was the weight wrong, or did the inputs change? By the end of Section 6 you will split the 3.2-point miss into its parts.

Exhibit 1: 0050 and 00679B: estimation window, test window and full sample
Full sample 2017-01 → 2026-10 Estimation window 2017–2021 Test window 2022-01 → 2026-10
Trading days (daily returns) 2,363 1,213 1,150
0050 total return +767.3% +138.0% +264.4%
00679B total return −17.3% +15.6% −28.5%
0050 annualized volatility 20.1% 16.4% 23.4%
00679B annualized volatility 13.5% 13.5% 13.4%
Correlation of daily returns −0.11 −0.22 −0.02
Minimum-variance weight in 0050, estimated in the window 32.7% 42.3% 25.3%
Lowest attainable volatility in the window 10.64% 9.22% 11.54%
Volatility of the 42.3% / 57.7% mix 10.91% 9.22% 12.44%
Source: FinMind TaiwanStockPrice and TaiwanStockDividendResult, retrieved 10 October 2026. Daily total returns between consecutive common trading dates from the base date 17 January 2017. Cash dividends are reinvested on every ex-date in the window (20 for 0050, 37 for 00679B), and 0050 is adjusted for its 4-for-1 split of 18 June 2025. Volatilities (divisor ) are annualized with trading sessions a year: the sample's 2,363 returns over 3,551/365.25 calendar years. Minimum-variance weights and volatilities use unrounded moments. Mixes are rebalanced daily. Computations: data/lm07_stats.py and companion notebook LM07_lab.ipynb.

Sections 2–5 extend derivatives, Taylor expansions and optimization to several variables and add equality constraints. Section 6 applies them to two assets and resolves the case, Section 7 adds inequality constraints and Section 8 solves everything in Python. Matrix notation for assets waits for LM10 and LM32.

NoteNotation for this module

A generic function of generic coordinates , instantiated per example, has partial derivatives and , gradient and Hessian . Constraints read or , with multipliers (equality) and (inequality) and optimal value . Bold letters are vectors: weights , exposures per unit of capital (Kelly fractions), expected returns , excess returns (), , a point , a direction , a step . For two assets, are annualized volatilities, the covariance, the correlation, the variance of ; is a portfolio variance, a volatility budget, a log growth rate, the smallest and largest curvatures of . Counts: returns, assets, iteration ; index assets or constraints. Moments use divisor , as in LM3; with 1,150 or more returns, would change volatilities by under 0.05%.

2Functions of Several Variables and Partial Derivatives

compute partial derivatives and total differentials of functions of two or three variables, and interpret them as sensitivities with the other inputs held fixed

A portfolio's volatility depends on two weights, two volatilities and a correlation at once, and a TWD investor's return on a US bond depends on the bond and on the currency.

Functions, surfaces and level curves

DefinitionFunction of two variables

A function of two variables assigns to each point of a domain one number . Its graph is the surface . A level curve (contour) is the set of points where takes a fixed value , .

Level curves are the map view of the surface. The central example is the variance of a two-asset portfolio. With weights and daily returns , the portfolio return deviates from its mean by . Squaring and averaging gives

The middle average is the covariance of the two returns, positive when they tend to fall on the same side of their means (LM16). Divided by the two standard deviations it gives the correlation . It lies between −1 and 1 by the Cauchy–Schwarz inequality , applied to the deviations (stated without proof; LM8). Annualizing each average with ,

The same expansion gives two facts used later. Averaging instead shows that the covariance of with the portfolio is , and weights give .

When , the level curves of are ellipses centred on the all-cash point , one for each variance level. The budget rule is a straight line across them. Section 5 shows that the minimum-variance portfolio sits where that line just touches an ellipse, and Exhibit 5 draws the picture.

Partial derivatives

DefinitionPartial derivative

The partial derivative of with respect to at is the ordinary derivative of the one-variable function :

is defined in the same way with held at .

To compute a partial derivative, treat every other variable as a constant and apply the rules of LM2. For the portfolio variance,

The first bracket is the covariance of asset 1 with the portfolio, found above. So the marginal variance of a position is twice its covariance with the portfolio it joins, a fact LM36 builds into risk budgeting.

Higher-order partial derivatives. Differentiating again gives , and the mixed partials and . For every function in this book the order does not matter.

TheoremSymmetry of mixed partial derivatives (Schwarz–Clairaut)

If and exist and are continuous near , then . Stated without proof; see Stewart, Clegg and Watson (2021, §14.3).

The total differential and the linear approximation

Move both inputs at once, by and . Going first along and then along , and using the one-variable linear approximation of LM2 on each leg, gives

When has continuous partial derivatives, the error is small relative to the step; when its second partial derivatives are continuous too, the error is of second order in . The linear part, , is the total differential. Dividing by a time step gives the multivariable chain rule: if and , then .

Example 1A US bond in Taiwan dollars: 00679B in 2022

A TWD investor's return on a US-dollar asset is , where is the asset's return in USD and the change in the USD/TWD rate. In 2022, 00679B returned −22.09% in TWD, while USD/TWD rose from 27.690 to 30.708, +10.90%. Back out the USD return implied for the fund, compute the partial derivatives there, and split the TWD return into bond, currency and cross terms.

Solution

The implied USD return solves : .

The partial derivatives are and , and . An extra 1% of bond return in USD adds 1.109% in TWD, because it is converted at a stronger dollar. An extra 1% rise in the dollar adds only 0.70%, because it applies to a bond position already down 30%.

Expanding the product gives the exact split . The linear approximation (7.1), , overstates the return by the cross term of 3.24 points: second order, but not small when the moves are 30% and 11%.

Example 2Marginal variance of a 60/40 mix

With the full-sample estimates for 0050 and 00679B, , and (so , , ), evaluate the variance, volatility and partial derivatives of at . Then estimate the change in volatility if (a) one percentage point of 0050 is added and financed with cash, or (b) one point is moved from 00679B to 0050.

Solution

, so . The partial derivatives are and .

Because , the chain rule gives : 0.182 for 0050 and 0.044 for 00679B.

(a) : percentage points (exact: 0.18). (b) : points (exact: 0.14). The answers differ because the experiments differ: (a) raises total exposure, while (b) stays on the budget line .

Pitfall"Holding everything else fixed" is not always possible

In a fully invested portfolio the weights cannot move one at a time: buying 0050 means selling something. So describes a trade financed with cash or margin, case (a) above; a budget-neutral switch needs the derivative along , case (b), or substitution of the constraint (Sections 5 and 6). Likewise, " holding volatilities fixed" ignores that in a crisis volatilities and correlations tend to move together.

導讀偏導數:凍結其他變數的敏感度

偏導數就是「其他條件不變」時的敏感度:只動一個變數,其餘全部凍結。台幣投資美債的報酬是 ,它對匯率的偏導數是 :2022 年債券本身已跌三成,美元再升 1%,台幣報酬只多 0.70%。全微分把各方向的敏感度加總,是一階近似;變動很大時(債跌三成、美元升一成),交叉項 就有 3.24 個百分點,不能忽略。在投資組合裡「其他條件不變」常常做不到:權重加總必須是 1,加碼 0050 就得減碼另一檔,要看的是沿著 方向的變化。

Knowledge Check 1

For at , , compute and . Use (7.1) to estimate the change in if rises a further 2 percentage points, to 7%, and explain why the estimate is exact here.

Answer

and . With and , (7.1) gives points. It is exact because is linear in when is fixed: the only second-order term, , vanishes when .

3Gradient, Hessian and the Multivariate Taylor Expansion

construct the gradient and Hessian of a function of several variables, and approximate the function with its second-order multivariate Taylor expansion

Directional derivatives and the gradient

Pick a direction, given by a unit vector , and walk from along it; the height of the surface along the walk is . By the chain rule, its slope at the start, the directional derivative, is

DefinitionGradient

The gradient of at is the vector of its partial derivatives, . Every directional derivative is a dot product with it: .

Two geometric facts follow. First, by the Cauchy–Schwarz inequality , with equality only when points along . So the gradient points in the direction of steepest ascent, and its length is the steepest slope.

Second, if a point moves along a level curve, then , and differentiating gives . So the gradient is perpendicular to the level curves.

For a step that is not a unit vector, is the first-order change in ; case (b) of Example 2 used exactly this, with .

The Hessian

DefinitionHessian

The Hessian of at is the matrix of its second partial derivatives,

symmetric when the Schwarz–Clairaut theorem applies.

The Hessian measures curvature in every direction at once. Differentiating twice with the chain rule gives

the curvature of along . For the portfolio variance the Hessian is constant, : twice the covariance matrix , the object LM10 studies.

The second-order Taylor expansion

The one-variable Taylor polynomial of LM3, applied to at , gives . Substituting the two derivatives just computed:

with all derivatives evaluated at . If has continuous third partial derivatives, the Lagrange remainder of LM3 (3.5), applied to , puts the error at for some : third order in . With three variables the pattern is the same, with three first derivatives and a Hessian. A polynomial of degree two is reproduced exactly by (7.2), and one of degree three once the third-order terms are added.

A Taylor attribution of the 42/58 mix's extra risk

Hold the weights of the case fixed at in 0050 and in 00679B. Their variance is then a function of three inputs,

Between the estimation and test windows the inputs moved from to . The partial derivatives at the starting point, from the unrounded estimates, are , and . Multiplying each by its input's change gives the first-order terms of Exhibit 2.

The first-order approximation explains 51.70 of the 69.67 %² rise in variance, or 74%. The Hessian terms supply the rest, slightly overshooting. The two largest are , with , and the cross term , with . A higher correlation hurts more when 0050 is more volatile, and only a second-order model can see that. Because is a polynomial of degree three in its inputs, the single third-order term , about −0.08, makes the expansion exact.

Exhibit 2: Taylor attribution of the variance of the 42.3% / 57.7% mix, 2017–2021 → 2022–2026
Input or term Change Partial derivative (%²)
(0050): 16.37% → 23.40% +0.070325 0.043994 +30.94
(00679B): 13.55% → 13.43% −0.001140 0.072465 −0.83
: −0.221 → −0.022 +0.199414 0.010827 +21.59
First-order approximation +51.70
Second-order approximation, adding +69.74
Exact change (degree-three polynomial) +69.67
Note: variance in squared percentage points (%²); it rose from 85.1 to 154.8, so volatility rose from 9.22% to 12.44%. Changes and derivatives in decimal units. Computed by data/lm07_stats.py.

To first order, 0050's higher volatility accounts for 44% of the rise in variance and the higher correlation for 31%. The second-order terms add 26%, half of it from the two moving together.

A two-variable Taylor model of portfolio growth

The most important use of (7.2) in this book is the growth rate of a levered portfolio. Hold a fraction of capital in 0050 and in 00679B, and finance the rest at an illustrative risk-free rate a year, a day. The daily return is , with excess returns . The quantity a Kelly investor maximizes (Kelly 1956; LM4) is the average log growth, annualized:

Differentiate under the sum: and . At these are the mean excess returns and minus the mean products, each divided by a power of .

On the data, , which is up to the factor . The Hessian has entries , against , and . The small gap in the first entry is the squared daily mean, smaller than the variance by a factor of order . Dropping it and the factors , (7.2) becomes

The same model follows from expanding day by day and averaging, as in LM3 (3.8). It neglects the squared daily mean, , and the terms of third and higher order in the daily portfolio return. In matrix form (7.3) reads , the starting point of LM43.

Exhibit 3: Exact sample log growth versus the quadratic model (7.3), 0050 and 00679B, 2017–2026
Portfolio (0050) (00679B) Exact Model (7.3) Model − exact
0050 alone 1.00 0.00 22.2200% 22.2297% +1.0 bp
00679B alone 0.00 1.00 −1.9569% −1.9573% 0.0 bp
Full-sample minimum-variance mix 0.33 0.67 6.6705% 6.6722% +0.2 bp
Twice 0050, financed at 2.00 0.00 38.8741% 38.9168% +4.3 bp
Half Kelly 2.80 −0.26 49.5495% 49.6964% +14.7 bp
Full Kelly of the model (Section 6) 5.59 −0.52 63.7424% 65.7619% +202.0 bp
Note: annual log growth; (illustrative); daily rebalancing, no costs. "Exact" is the sample average of times . Growth is evaluated at unrounded exposures.

The model is excellent for unlevered mixes and degrades as leverage grows. For 0050 alone the 1.0 bp error is the squared mean's 1.2 bp less 0.2 bp of higher-order terms. At full Kelly the model overstates growth by 2.02 points a year: 0.35 from the dropped squared daily mean and 1.67 from the third- and higher-order terms of , which grow with the daily moves. The fourth-order term alone costs 1.87 points; the third adds back 0.49, and fifth and higher orders cost 0.29.

導讀梯度、海森矩陣與多變數泰勒展開

梯度 指向函數上升最快的方向,長度就是最陡的坡度,而且永遠和等高線垂直。海森矩陣 描述各個方向的彎曲程度。多變數泰勒展開把 LM3 的「高度+斜率+曲率」推廣到多個變數:。拆解 42/58 組合的變異數為何上升:一階項只解釋 74%,加上二階項(尤其是 0050 波動與相關係數「一起」上升的交叉項)幾乎完全吻合。同一個展開也給出凱利成長率的二次近似:槓桿越高越不準,滿凱利時每年高估約 2 個百分點,其中 0.35 個百分點來自被捨去的日均報酬平方,其餘來自三階以上的高階項。

Knowledge Check 2

At the growth function has gradient . In which unit direction does rise fastest, and at what rate? What is the rate of change along the budget-neutral unit direction ?

Answer

Steepest ascent is along , at rate . Along the rate is , smaller, as Cauchy–Schwarz requires.

4Unconstrained Optimization in Two Variables

locate and classify the critical points of a function of two variables with the first-order conditions and the 2×2 definiteness test

LM4 found optima of one-variable functions with two conditions: a zero derivative and the sign of the second derivative. Both carry over to two variables, the second with more care, because a surface can curve up in one direction and down in another.

First-order conditions

If is differentiable and has a local maximum or minimum at an interior point , then every one-variable slice through has zero slope there. In particular , that is, . Points where the gradient vanishes are critical (or stationary) points. As in one variable, the condition is necessary, not sufficient.

Second-order conditions: the 2×2 test

At a critical point the linear term of (7.2) vanishes, so for small steps , and the sign of this quadratic form decides the shape. If , completing the square gives

If , both terms have the sign of , and the form is non-zero for every . If , the two terms have opposite signs, and the form takes both signs. When has continuous second partial derivatives, the remainder of (7.2) is small next to near and cannot overturn a definite form (stated without proof; Stewart, Clegg and Watson 2021, §14.7). When the test is silent and higher-order terms decide: has a minimum at the origin and does not, with the same Hessian there. Exhibit 4 collects the cases. In matrix language, is positive definite exactly when and , the case of a criterion LM10 extends to assets.

Exhibit 4: The second-derivative test at a critical point of
Condition Quadratic form Shape Example in this module
and positive definite local minimum (bowl) portfolio variance with
and negative definite local maximum (dome) the growth model (7.3)
indefinite saddle point at the origin (Example 3)
semidefinite or degenerate test inconclusive with : a valley floor
Note: at the critical point; the test classifies local behaviour only.

Global optima and convexity

A local test says nothing about the rest of the surface. One condition does: if for every at every point (the Hessian is positive semidefinite everywhere), is convex, and any critical point is a global minimum (stated without proof; LM34 develops convexity). Quadratic functions have a constant Hessian, so for them the local test is global.

Portfolio variance has with : it is convex, and strictly convex when . The growth model (7.3) has and is concave.

Example 3A saddle and a local minimum

Find and classify the critical points of .

Solution

gives , and gives . Hence , so or , and the critical points are and . The second derivatives are , and .

At : , a saddle. Indeed while for small .

At : and , a local minimum with . It is not global, because as . Without convexity, a local test cannot certify a global optimum.

PitfallWhen : perfectly correlated assets

With the variance Hessian has . The minimum is then a whole line of weights (a valley floor) or a riskless long–short hedge. The growth model (7.3) is then unbounded unless (for ) or (for ), where is asset 's annualized Sharpe ratio. Otherwise a riskless spread with a positive excess return can be levered without limit.

Real assets are never exactly collinear, but correlations near ±1 make tiny. Optimizers then return huge, unstable long–short weights; LM9 and LM11 treat this ill-conditioning.

Knowledge Check 3

Classify the critical point of . Every coefficient is positive; is the origin a minimum?

Answer

, , , so : a saddle, not a minimum. For example . Positive coefficients do not make a quadratic form positive definite; the cross term must satisfy .

5Equality Constraints: Lagrange Multipliers

solve equality-constrained problems with Lagrange multipliers, and interpret the multiplier as a shadow price

Most portfolio problems come with constraints: the weights must sum to one, the volatility must hit a target, an order must be completed. When a constraint can be solved for one variable, substitution reduces the problem to LM4. Lagrange's method works without solving the constraint, and it delivers a bonus: the price of the constraint.

The tangency condition

Consider optimizing subject to , and walk along the constraint curve. At a constrained optimum the slope of along the curve must be zero, so is perpendicular to the curve there. The constraint curve is a level curve of , so is perpendicular to it too (Section 3). Two vectors perpendicular to the same curve are parallel: at the optimum, the level curve of just touches the constraint curve. Exhibit 5 shows this for the minimum-variance portfolio.

Algebraically, suppose at the optimum. Near it the constraint defines as a function of (the implicit function theorem, stated without proof), and differentiating gives , so . The one-variable function has a critical point at the optimum:

Writing , this says , and holds by definition of .

TheoremLagrange multiplier condition

Let and have continuous partial derivatives, and let be a local maximum or minimum of subject to , with . Then there is a number with

Equivalently, is a critical point of the Lagrangian .

The three conditions give candidates; whether a candidate is a maximum, a minimum or neither needs a second-order or a global argument. A convex objective (for a minimum) or a concave one (for a maximum) with a linear constraint makes a candidate the global answer, as in the minimum-variance problems of Section 6; Section 7 states the general result. A curved constraint can produce several candidates: Example 5 has two.

Exhibit 5: Level curves of portfolio variance and the budget line, 2017–2021 estimates
Note: ellipses are sets of weights with equal volatility; the dashed line is the budget . The smallest ellipse that meets the line touches it at the minimum-variance weights (42.3%, 57.7%). There , perpendicular to the ellipse, is parallel to , perpendicular to the line.

The multiplier is a shadow price

ResultShadow price (envelope theorem)

Let be the optimal value when the constraint level is , and suppose the solution moves smoothly with . Then

Proof. By the chain rule and (7.5), . Differentiating the identity shows that the bracket equals 1.

This is the envelope theorem of LM4 (4.16) applied to the Lagrangian, whose partial derivative in is . So , with an error of second order in . The units of are units of objective per unit of constraint: NT$ per contract, return per unit of risk, variance per unit of capital.

Example 4Splitting a TX order between two sessions

A trader must buy TX contracts, in the day session and after hours. Model the temporary-impact cost as with illustrative coefficients and (NT$ per contract squared), the night book being thinner. Find the cost-minimizing split and interpret the multiplier.

Solution

. The conditions , and give , so , and , in NT$ per contract.

The multiplier is the marginal cost of the last contract, and the optimum equalizes it across sessions: if one session were cheaper at the margin, moving a contract there would cut the total.

Shadow price: the optimal cost is , so (NT$) and , as (7.6) requires. The 101st contract adds NT$287.1 exactly; the difference from is the second-order term .

In whole contracts a 71/29 split costs NT$14,287, essentially optimal; an equal split costs NT$17,500, 22.5% more.

Example 5A risk budget across two strategies

A desk runs a TX trend strategy (expected excess return , volatility ) and a futures–spot basis strategy (, ), with correlation — illustrative figures. Exposures per unit of capital can be levered. Maximize expected excess return subject to a volatility budget , and interpret the multiplier.

Solution

Maximize subject to , with . The Lagrangian conditions are

This is a linear system in ; eliminating one unknown (Cramer's rule, LM9) gives and , a ratio of 1 to 7.43. The constraint is an ellipse, so scaling to 10% volatility gives two candidates, , with . The minus sign is the minimum, an excess return of . The maximum is , with : the low-volatility basis trade is levered about two times. Posed as , the problem rules out the second candidate by the sign condition of Section 7.

The expected excess return is 8.31%, against 5.0% for the trend strategy alone at 10% volatility and 7.5% for the basis strategy alone. It equals , where is the best attainable Sharpe ratio, derived in Section 6.

The multiplier on the variance budget is . By (7.6), the optimal excess return rises by 4.15 per unit of variance budget. In volatility units , so raising the budget from 10% to 11% adds 0.83 percentage points. The shadow price of risk is the best attainable Sharpe ratio.

導讀拉格朗日乘數:限制式的影子價格

在限制條件下求極值,最適點一定落在目標函數的等高線與限制曲線「相切」之處,此時兩者的梯度平行:。乘數 不只是解題工具,它是「影子價格」:把限制放寬一單位,最適值大約改變 ,誤差是二階的。拆單例子裡, 是最後一口台指期的邊際成本,日盤與夜盤的邊際成本必須相等,否則把一口移到便宜的時段就能省錢。風險預算例子裡,以波動度計的乘數就是可達到的最高夏普比率。看到乘數,先問它的單位,再問放寬限制值不值得。

Knowledge Check 4

In Example 4, night-session liquidity improves to (NT$ per contract squared). Find the new split and multiplier, and the approximate cost of a 101st contract.

Answer

Now with : , and , in NT$ per contract. The 101st contract costs about NT$240 (exactly , since ).

6Two-Asset Portfolios by Calculus

derive the two-asset minimum-variance weight and explain how correlation determines the diversification benefit
derive the two-asset tangency weights and Kelly fractions, and relate them to

This section applies Sections 2–5 to two assets, in the tradition of Markowitz (1952) and of Merton's (1972) closed forms: the minimum-variance weight, in and out of sample, then the tangency portfolio and the Kelly fractions.

Minimum variance, by substitution and by Lagrange

Substitution. With in asset 1 and in asset 2, the variance is a one-variable function,

Its derivative is . Setting it to zero:

The second derivative is , where (Section 2). It is positive unless the two returns differ by a constant, so is the global minimum. Substituting back:

Because is a quadratic in with second derivative , it can be written exactly around its minimum:

Lagrange. Minimize subject to . With , the conditions are

Subtracting gives , which with is (7.7) again. The two conditions say more: at the minimum, each asset has the same covariance with the portfolio.

Multiply the first condition by , the second by and add: , so . That is the shadow price (7.6) of the budget: investing units in the same proportions gives variance , whose derivative at is .

Correlation and diversification

Formula (7.8) shows how correlation sets the benefit of diversification.

  • . Then only through a short position: when . Long-only, volatility is linear in , and there is no diversification.
  • . Then : each asset is weighted by its inverse variance.
  • . Then and : a perfect hedge without shorting.

The minimum-variance portfolio shorts the riskier asset when the numerator of (7.7) turns negative, , that is, when (asset 1 being the riskier one). For 0050 and 00679B the ratio is , far above any correlation in the sample, so the minimum is interior.

The envelope theorem of LM4 (4.16) prices a change in without differentiating (7.8): because , at . With the 2017–2021 estimates this is 0.0108, so a correlation 0.1 higher raises by 0.00108 (exact to that precision) and from 9.22% to 9.79%.

Back to the case: in sample and out of sample

Example 6The minimum-variance mix estimated at the end of 2021

Over 2017–2021 the annualized estimates are (0050), (00679B) and , so , and . Compute , and .

Solution

The numerator of (7.7) is and the denominator , so in 0050 and in 00679B.

From (7.8), , so , and .

The in-sample calculation is correct: the 42/58 mix's realized volatility over 2017–2021 is exactly 9.22%, because the sample variance of a fixed mix is at the sample moments (Section 2).

The calculation was right; the inputs changed, as Exhibit 6 shows year by year. In 2022 both funds fell by about 21–22% while their daily correlation was −0.06: a correlation near zero does not prevent a joint loss in a year when both assets drift down.

From 2023 the trailing one-year correlation moved between −0.29 (April 2025) and +0.25 (December 2023) (Exhibit 7). At lower frequencies the correlation turned positive: between the two windows it rose from −0.29 to +0.21 for monthly returns and from −0.15 to +0.09 for weekly ones.

Exhibit 6: 0050 and 00679B by calendar year
Year Days 0050 return 00679B return 0050 volatility 00679B volatility Correlation Minimum-variance weight in 0050
2017 (from 17 Jan) 235 +17.3% −0.2% 9.3% 8.4% −0.04 45.2%
2018 247 −5.0% +0.7% 16.3% 9.0% −0.39 30.0%
2019 242 +33.5% +13.7% 12.2% 12.2% −0.43 50.1%
2020 245 +31.1% +8.1% 22.8% 21.3% −0.17 47.1%
2021 244 +22.0% −6.4% 17.6% 12.5% −0.17 36.1%
2022 246 −21.3% −22.1% 21.6% 16.1% −0.06 36.4%
2023 239 +27.4% +3.1% 13.9% 12.6% +0.20 43.8%
2024 242 +48.7% −2.9% 24.6% 11.0% −0.24 21.8%
2025 238 +36.9% −0.6% 24.5% 15.7% +0.01 28.9%
2026 (to 8 Oct) 185 +78.7% −7.7% 30.6% 9.6% +0.06 7.5%
Note: days are common trading dates (2025 omits 11–17 June, when 0050 was halted for its split); total returns with dividends reinvested; volatilities annualized with ; correlations of daily returns; the last column applies (7.7) to each year's own unrounded estimates. Source: FinMind, retrieved 10 October 2026; data/lm07_stats.py.
Exhibit 7: Trailing one-year correlation of daily returns, 0050 and 00679B
Note: correlation over the trailing 243 trading days; the shaded band is calendar 2022. The dashed line marks the correlation over each full window, −0.22 for 2017–2021 and −0.02 for 2022–2026.

Identity (7.9) now splits the miss exactly, and Exhibit 8 draws it. With the test-window inputs, the lowest attainable variance was %² at , and the coefficient of the quadratic term was , half the curvature . The in-sample weight was 0.17067 (17.07 points) too high, which adds %². Together, %², the 12.44% volatility the 42/58 mix actually realized.

ResultResolution of the case

The 3.2-point volatility miss had two sources. The minimum itself moved, from 9.2% to 11.5%, because 0050 became more volatile and the correlation rose toward zero. The weight error added only 0.9 points, because by (7.9) the penalty is quadratic in the weight error: on the test-window curve, a 5-point error costs 0.08 points of volatility and a 10-point error 0.32. The realized 12.44% was still below 00679B's 13.43% and far below 0050's 23.40%, so diversification survived, only smaller than estimated.

Exhibit 8: Portfolio volatility against the weight in 0050, estimated in and out of sample
Note: volatility of a daily-rebalanced mix from (7.8)–(7.9) with each window's moments. A: in-sample minimum, 9.22% at . B: ex-post minimum, 11.54% at . C: the in-sample weight in the test window, 12.44%.

Minimum variance ignores expected returns, which mattered here: over the test window's 4.77 calendar years the 42/58 mix compounded at 8.67% a year and 0050 at 31.12%, while 00679B lost 6.78% a year.

The tangency portfolio

The portfolio with the highest Sharpe ratio, , is the tangency portfolio: on a chart of expected return against volatility, it is where a line from the risk-free rate touches the set of risky portfolios (LM33).

Example 5 already solved this problem. Maximizing expected excess return for a fixed volatility gave exposures proportional to , whatever the budget. Scaling exposures by a positive constant leaves the Sharpe ratio unchanged, so the same direction maximizes it. Normalizing it to sum to one gives the tangency weights

The highest Sharpe ratio follows by evaluating the ratio on that direction. With , it is . Dividing numerator and denominator by :

Subtracting from (7.11) leaves . So adding asset 2 raises the best Sharpe ratio unless , which is exactly when its weight in (7.10) is zero. In Example 5, , so .

When the denominator of (7.10) is near zero the tangency weights explode, as the test-window estimates in Exhibit 9 show.

The two-asset Kelly fractions

The Kelly investor maximizes the growth rate. With the quadratic model (7.3), the first-order conditions are

a linear system. Eliminating one unknown (Cramer's rule, LM9) gives

The Hessian of (7.3) is , with and whenever : by Exhibit 4, a maximum, and global because (7.3) is concave. Three readings of (7.12) connect it to the rest of the book.

  1. Matrix form. A matrix times a vector is the vector of each row's dot product with it. The matrix undoes multiplication by (LM9), and is exactly (7.12). So , the multi-asset Kelly rule of LM43.
  2. Kelly is tangency, levered. The numerators of (7.12) and (7.10) are the same. Hence , with total exposure . The Kelly investor holds the tangency portfolio and chooses only how much of it to hold.
  3. Growth at the optimum. At the conditions give , so (7.3) becomes . With one asset, : LM4 derived with no financing cost, and LM42 adds .
Example 7Kelly fractions for 0050 and 00679B, full sample

Use the full-sample estimates of Example 2, arithmetic means and , and an illustrative , so and . Compute the Kelly fractions (7.12), the tangency weights, and the growth rate of the quadratic model, and compare with the exact sample growth.

Solution

. The numerators are and . So and : in-sample, the Kelly investor would borrow to hold 5.6 times capital in 0050 and short half a unit of 00679B.

Total exposure is . The tangency weights (7.10) divide each numerator by their sum, 0.0036933: . With , and , (7.11) gives , hardly above 0050's own 1.132: in this sample the bond adds almost nothing to the best Sharpe ratio.

The model's growth rate is a year. The exact sample growth at is lower, 63.7% (Exhibit 3). The exact optimum, found numerically in Section 8, is with 64.0%. On 7 April 2025 alone, the position at would have lost 56.9% of capital, and the exact optimum 53.7%.

Exhibit 9: Kelly fractions and tangency weights estimated on three windows
Window (0050) (00679B) Tangency weights
2017–2021 17.2% +2.3% 16.4% 13.5% −0.22 7.00 +3.14 10.14 69% / +31% 1.13
2022–2026 28.6% −7.7% 23.4% 13.4% −0.02 5.17 −4.06 1.11 466% / −366% 1.34
2017–2026 22.8% −2.5% 20.1% 13.5% −0.11 5.59 −0.52 5.07 110% / −10% 1.13
Note: are arithmetic annualized means minus an illustrative ; fractions from (7.12), tangency weights from (7.10) and from (7.11), all computed from unrounded estimates. In-sample estimates, shown to illustrate instability, not as recommendations.
PitfallOptimal weights fitted to the past are not recommendations

Exhibit 9 shows how fragile the inputs are. The 00679B fraction flips from 3.1 times long to 4.1 times short between windows, and the tangency weights go from a sensible 69/31 to an absurd 466/−366. Expected returns are the culprit. The standard error of a mean annual return, its typical estimation error, is about (LM20). Over the 4.95 calendar years of 2017–2021 that is 7.36 points for 0050, 39% of its 18.7% mean, and 6.09 points for 00679B, more than its 3.8% mean.

Chopra and Ziemba (1993) found errors in means far more damaging to optimal portfolios than errors in variances or covariances. Minimum-variance weights avoid means altogether, yet Exhibit 6 shows that variances and correlations move too. LM35 (estimation error and shrinkage) and LM45 (Kelly under uncertainty) take up the repair.

導讀最小變異數、切點組合與凱利比例

兩資產最小變異數權重 只用到波動度與共變異數,不需要預期報酬。用 2017–2021 的資料算出 0050 占 42.3%、預測波動 9.22%;2022–2026 實際卻是 12.44%。差距主要來自「輸入值變了」:0050 波動從 16.4% 升到 23.4%,相關係數從 −0.22 升到 −0.02(月報酬更從 −0.29 翻正為 +0.21);權重估錯 17 個百分點只多付 0.9 個百分點的波動,因為變異數曲線在最低點附近很平坦。切點組合與凱利比例還需要預期報酬,估計誤差大得多:00679B 的凱利比例在不同樣本期間從 +3.14 變成 −4.06。凱利就是「加了槓桿的切點組合」,但樣本內算出的最適解不是投資建議。

Knowledge Check 5

Two assets have the same volatility and correlation . Show that the minimum-variance weights are 50/50 and that . What happens as and as ?

Answer

With and , (7.7) gives . Then . As the risk vanishes (a perfect hedge); as it tends to (no diversification). At the volatility falls by the factor .

7Inequality Constraints: The KKT Conditions

apply the Karush–Kuhn–Tucker conditions to long-only, policy and leverage constraints, and identify which constraints bind

Real mandates are full of inequalities: no short sales, at least 60% in equities, exposure at most twice capital. LM4 handled one inequality in one variable: a maximum on an interval is either interior, with , or at an endpoint where points outward. The Karush–Kuhn–Tucker conditions, found by Karush (1939) and again by Kuhn and Tucker (1951), extend that logic to several variables and constraints; Kjeldsen (2000) tells the history, and LM34 applies the conditions to larger portfolios.

Slack or binding

At the optimum each inequality constraint is either slack, and could be deleted without changing the answer, or binding, and behaves like an equality with a multiplier. The multiplier's sign must say that the constraint pushes in the right direction: for a maximization with constraints , relaxing a constraint can only help, so its shadow price cannot be negative.

TheoremKarush–Kuhn–Tucker (KKT) conditions

Let and every have continuous partial derivatives. Maximize subject to , one inequality for each . If is a local maximum and the gradients of the binding constraints are linearly independent (for two of them: not parallel), there are multipliers with

If is concave and every is convex, these conditions are also sufficient for a global maximum. Stated without proof; see Nocedal and Wright (2006, Theorem 12.1) for necessity and Boyd and Vandenberghe (2004, §5.5.3) for sufficiency. For a minimization, apply (7.13) to . An equality constraint enters as in Section 5, with a multiplier of either sign, and must be linear for sufficiency.

The four parts are stationarity, feasibility, dual feasibility and complementary slackness. Stationarity says that the uphill direction is a non-negative combination of the binding constraints' outward normals, so every direction that would raise leaves the feasible set. Complementary slackness says that each constraint either binds or has a zero multiplier. As in (7.6), is the rate at which the optimal value rises when is relaxed.

The practical method is enumeration. Guess which constraints bind, solve stationarity together with those constraints as equalities, and keep the candidate only if it is feasible and every multiplier is non-negative. With two variables and three constraints there are at most eight binding sets to try.

Example 8Kelly with a leverage cap and no short sales

Maximize the growth model (7.3) with the full-sample estimates of Example 7 subject to , and . The first constraint, illustrative, caps net (total) exposure at twice capital; it equals gross exposure only when both positions are long. Identify the binding constraints and interpret the multipliers.

Solution

The unconstrained optimum violates the cap and the no-short rule. Of the eight possible binding sets, the one with all three constraints is impossible ( breaks ), which leaves seven; Exhibit 10 solves each. Only one passes every test: the cap and bind, at .

Write the constraints as , and , with multipliers , and . Stationarity reads and , with because is slack at .

At , , so . Next, , so . Both multipliers are non-negative.

In-sample, at the margin each unit of exposure is worth 14.7 points of model growth a year, the figure to compare with a margin lender's spread over . That is a first-order value: exactly, raising the cap to 2.1 adds 1.45 points, and to 3 adds 12.64.

needs care. Relaxing to lets the investor short 1% of 00679B and, because the cap limits net exposure, hold 2.01 in 0050. Re-solving gives and 0.166 points more growth, as predicts. The extra 0050 adds 0.1465 and the short 0.0196: since , selling 00679B raises growth. Under a gross cap the short would use up cap instead: loses 0.127 points, and stays optimal.

With the exact growth function instead of (7.3), the optimum is the same and the cap's multiplier is 0.146 (Section 8).

Exhibit 10: KKT candidates for the capped, long-only Kelly problem, full sample
Binding constraints Candidate Multipliers Feasible All multipliers ≥ 0 Model growth
none (5.59, −0.52) — no: breaks the cap, shorts 00679B yes (none) 65.76%
cap only (4.58, −2.58) no: shorts 00679B yes 60.42%
only (5.63, 0.00) no: breaks the cap yes 65.52%
only (0.00, −1.40) no: shorts 00679B no 3.28%
cap and (2.00, 0.00) , yes yes 38.92%
cap and (0.00, 2.00) , yes no −7.24%
and (0.00, 0.00) , yes no 1.50%
Note: model (7.3) with full-sample estimates and (illustrative). The cap is on net exposure ; "cap only" has gross exposure . The optimum is the only candidate that is feasible with non-negative multipliers.

The same logic covers long-only minimum variance. When , (7.7) puts a negative weight on the riskier asset; with no short sales the optimum is the corner that holds only the safer asset, and the multiplier is the slope of there. The widget in Section 6 reports it, and Practice Problem 9 works an example.

PitfallChecking only some of the KKT conditions

Every candidate in Exhibit 10 satisfies stationarity. Four are infeasible, and three of those beat the optimum's growth because they ignore a constraint. Two feasible candidates, and , carry a negative multiplier: relaxing that constraint would lower growth, so it is not what holds the solution back. Feasibility alone leaves three candidates and the sign test alone four; only passes both. Check all four conditions of (7.13), and state which constraints bind: the binding set is what changes when the mandate changes.

導讀KKT 條件:限制式是緊還是鬆

不等式限制(不可放空、最低股票比重、曝險上限)在最適點只有兩種狀態:沒碰到(鬆弛),就當它不存在,乘數為 0;碰到了(緊),就當成等式處理,乘數必須 ≥ 0,表示這個限制確實擋住了你想去的方向。這就是互補鬆弛性:乘數與鬆弛量至少有一個是 0。實務解法是列舉哪些限制是緊的,對每個候選解檢查四個條件,缺一不可。槓桿上限的乘數是邊際上「多一單位曝險能多賺多少成長率」,樣本內約每年 14.7 個百分點,可以直接拿來和融資利差比較;這是一階估計,上限從 2 放寬到 3 實際多 12.64 個百分點。要注意例子裡的上限管的是淨曝險 :放空 00679B 會騰出額度加碼 0050;若上限管的是總曝險 ,放空反而會占用額度。

Knowledge Check 6

An investment policy, illustrative here, requires at least 60% in 0050. With the full-sample estimates of Example 2, does this floor bind for the minimum-variance mix? If so, find its multiplier and the cost of the policy.

Answer

Yes: (7.7) gives , so the floor binds at . For the maximization of , with on the budget and on , stationarity reads and . Hence , the slope of along the switch direction of Example 2. A 59% floor would cut the variance by about (exactly 0.000344). Against the unconstrained minimum of 10.64%, the policy costs 2.05 points of volatility (12.69%).

8Multivariable Optimization in Python

implement numerical gradients, Newton's method and constrained optimization in Python, and distinguish in-sample optimal weights from out-of-sample results

Closed forms such as (7.7) and (7.12) need quadratic problems in two variables; the exact growth function is not quadratic, and real problems have more assets and constraints. Three numerical tools cover most cases: finite-difference derivatives, Newton's method and a constrained solver.

Numerical gradients and Hessians

Central differences carry over from LM2 coordinate by coordinate, with an error of order for a step . For a mixed partial derivative, difference in both coordinates (the lab applies this to the Hessian):

The code below builds the exact growth function of Section 3 from the lab's daily returns and checks its gradient at against the mean excess returns.

Python
import numpy as np

# from Part 1 of the lab: daily total returns R1, R2 (0050, 00679B)
# and A = T/years ≈ 243.0543
rf = 0.015                                    # illustrative r_f
rd = rf / A                                   # daily financing rate r_f/A
X = np.column_stack([R1, R2]) - rd            # daily excess returns, T x 2

def g(f):
    """Annualized sample log growth of exposures f, financed at r_f."""
    return A * np.log1p(rd + X @ f).mean()

def num_grad(fun, x, d=1e-5):
    e = np.eye(len(x))
    return np.array([(fun(x + d * e[i]) - fun(x - d * e[i])) / (2 * d)
                     for i in range(len(x))])

print(num_grad(g, np.zeros(2)).round(4))    # [ 0.2275 -0.0255]  ~ (e1, e2)

Gradient ascent and Newton's method

Gradient ascent repeats : step uphill, re-measure the slope, repeat. For the quadratic model (7.3), whose Hessian is , let and be the smallest and largest curvatures over unit directions: the roots of , the eigenvalues of (LM11). The best fixed step is , and it shrinks the distance to the optimum at least by the factor per step, where is the condition number (stated without proof; see Nocedal and Wright 2006). For the full-sample , and , so , and the factor is 0.39.

Near its optimum the exact growth function is more curved, with curvatures 0.0198 and 0.0493. The same step therefore contracts the distance by per step, and the growth shortfall by about .

Newton's method maximizes the second-order Taylor model (7.2) of at the current point. Setting the model's gradient with respect to the step to zero gives the linear system , solved as in Section 6, and : the two-variable version of LM4's iteration. On a quadratic it lands on the optimum in one step. On the exact growth function, the first step from lands at , near the model optimum , because at the Taylor model is essentially (7.3). Two more steps reach .

Python
def grad_hess(f):
    """Exact gradient and Hessian of g:
    A*mean(X/(1+R_p)) and -A*mean(X X'/(1+R_p)^2)."""
    w = 1 / (1 + rd + X @ f)
    return A * (X * w[:, None]).mean(0), -A * (X.T * w**2) @ X / len(X)

f = np.zeros(2)
for k in range(1, 5):
    gr, H = grad_hess(f)
    f = f + np.linalg.solve(H, -gr)    # Newton step: H Delta = -grad g
    print(k, f.round(4), f"{g(f):.4%}")
# 1 [ 5.5621 -0.5156] 63.7851%
# 2 [ 5.3231 -0.2675] 64.0296%
# 3 [ 5.3186 -0.2641] 64.0296%
# 4 [ 5.3186 -0.2641] 64.0296%
Exhibit 11: Growth shortfall by iteration: gradient ascent versus Newton's method
Note: shortfall of the exact sample growth from its maximum, 64.03% a year, starting at ; gradient ascent uses the fixed step . Each series ends at its first shortfall below , drawn on the bottom edge: far smaller shortfalls approach the rounding error of the computed growth rate (about ) and vary from machine to machine.

Newton needs two iterations to bring the shortfall below and three to go below ; gradient ascent needs 18 and 30 (Exhibit 11). Newton's shortfall falls from to and then below : the number of correct digits roughly doubles at each step. The price is the Hessian: second derivatives and a linear solve per step, cheap for two assets and expensive for two thousand. Gradient methods dominate large problems for that reason, though they struggle when runs into the thousands.

Constrained problems with scipy.optimize.minimize

For constraints, scipy.optimize.minimize with method="SLSQP" (sequential least-squares quadratic programming; Kraft 1988) handles bounds, equalities and inequalities. It solves a sequence of quadratic models, each a small KKT problem. Verify a multiplier the way (7.6) defines it: re-solve with the constraint relaxed and divide the change in the optimum by the change in the constraint. Recent SciPy versions also report the multipliers in the result object.

Python
from scipy.optimize import minimize

def kelly_capped(cap):
    """Maximize exact sample growth subject to f1 + f2 <= cap and f >= 0."""
    res = minimize(lambda f: -g(f), x0=[0.5, 0.5], method="SLSQP",
                   bounds=[(0, None), (0, None)],
                   constraints=[{"type": "ineq",
                                 "fun": lambda f: cap - f.sum()}])
    return res.x, -res.fun

f_cap, g_cap = kelly_capped(2.0)
print(f_cap.round(4), f"{g_cap:.4%}")    # [2. 0.] 38.8741%
nu_cap = (kelly_capped(2.01)[1] - kelly_capped(1.99)[1]) / 0.02
print(f"shadow price of the cap: {nu_cap:.4f}")    # 0.1461

The exact solution agrees with Example 8: the cap and the no-short rule bind, and the cap's shadow price is 0.146 against the model's 0.147.

In PythonCompanion lab

The Polars-first notebook LM07_lab.ipynb builds the total-return series, with all 37 00679B ex-dates and the 0050 split. It reproduces the data exhibits and the examples, and solves the computational practice problems.

The notebook reads FinMind files that you download once with your own token. From the repository root, run uv run python data/fetch_finmind.py, then uv run python data/make_lab_data.py. FinMind's licence does not allow the book to redistribute the data. After that, the notebook runs offline.

Knowledge Check 7

Why does Newton's method reach the maximum of the model (7.3) in one step from any starting point, but need three steps on the exact growth function?

Answer

Each Newton step maximizes the second-order Taylor model at the current point. For (7.3) that model is the function, because (7.3) is quadratic, so the first step lands on . The exact growth function has third- and higher-order terms, so each step lands only near the optimum, with a new error of second order in the old one.

Summary

  • Partial derivatives are sensitivities with the other inputs frozen; the total differential (7.1) adds them to first order. For 00679B in 2022 it missed a cross term of −3.24 points.
  • The gradient points uphill, perpendicular to level curves; the Hessian measures curvature. The second-order expansion (7.2) explained nearly all of the rise in the 42/58 mix's variance, first-order terms 74%.
  • The quadratic growth model (7.3) is accurate to about 1 bp unlevered. At full Kelly it overstates growth by 2.02 points: 0.35 from the dropped squared daily mean, 1.67 from higher-order terms.
  • Interior optima have ; the test (7.4) classifies them through and , and convexity makes a local minimum global.
  • At a constrained optimum (7.5), and the multiplier is the shadow price (7.6): a marginal cost, or the best Sharpe ratio for a volatility budget.
  • The minimum-variance weight (7.7) needs no expected returns, and . Fitted on 2017–2021 it predicted 9.22% volatility; 2022–2026 delivered 12.44%, mostly because the inputs changed. By (7.9), the 17-point weight error cost only 0.9 points.
  • Kelly fractions (7.12) are the tangency portfolio (7.10) levered by , with from (7.11). Their estimates swing violently between windows: in-sample optima are not recommendations.
  • The KKT conditions (7.13) add non-negative multipliers and complementary slackness. Enumerating binding sets solves small problems, Newton's method and SLSQP larger ones, and bumping a constraint checks its multiplier.

Practice Problems

  1. For , the mixed partial derivative at is:

    • A.4
    • B.14
    • C.16
  2. A TWD investor holds a USD bond fund. Over a quarter the fund returns −5% in USD and USD/TWD rises 3%. Compute the TWD return exactly, its linear approximation and the cross term. Then use to estimate the TWD return had USD/TWD risen 4% instead, and compare with the exact value.

  3. The two-period geometric mean return is . Derive its second-order Taylor expansion about , evaluate it at , , and compare with the exact value. Relate the quadratic term to the volatility drag of LM3.

  4. Find and classify all critical points of .

  5. At the minimum-variance portfolio of two assets the budget multiplier is . If the investor scales the budget from 1 to 1.05, keeping the proportions, the minimum variance rises by approximately:

    • A.0.0010
    • B.0.0020
    • C.0.0005
  6. Asset 1 has volatility 30% and asset 2 has 15%. For which correlation does the unconstrained minimum-variance portfolio short asset 1?

    • A.0.30
    • B.0.45
    • C.0.60
  7. Two assets have expected excess returns and , volatilities and , and correlation . Compute the tangency weights and the tangency Sharpe ratio, and compare it with each asset's Sharpe ratio.

  8. For the assets of Problem 7, compute the Kelly fractions from the growth model (7.3) and the total exposure. Verify that , and compute the model growth rate with .

  9. A broad index fund has and a leveraged fund on a related index has , with (illustrative). (a) Find the unconstrained minimum-variance weights and volatility. (b) Find the long-only optimum and its KKT multiplier. (c) Interpret the multiplier.

  10. Maximize the growth model (7.3) for , , , , , subject to , , . Report the binding constraints and the multipliers, and compare the excess growth with the unconstrained optimum.

  11. Which statement about in-sample optimization is correct?

    • A.The in-sample minimum-variance weight minimizes future volatility.
    • B.Near the minimum, a weight error raises variance in proportion to the square of the error.
    • C.Kelly fractions estimated from ten years of daily data are accurate to within ±0.1.
  12. (Python) With the lab data, estimate the minimum-variance weight (7.7) on each calendar year from 2017 to 2025 and hold that weight, rebalanced daily, through the following year. In how many of the nine following years did it produce lower volatility than a 50/50 mix, also rebalanced daily? Compare the simple averages of the nine annual volatilities with those of the 50/50 mix and of each year's own (ex-post) minimum-variance weight.

Solutions

  1. B is correct. , so ; as well. A is , and C is , the first derivative rather than the mixed one.

  2. Exactly, . The linear approximation is and the cross term . With , an extra 1 point of currency gain adds 0.95 points, giving ; exactly, . The estimate is exact because is linear in for fixed .

  3. At : , , and . By (7.2), . At : , against the exact , an error of 4.8 bp from third- and higher-order terms. The two returns have mean and population standard deviation , so the quadratic term is exactly the volatility drag of LM3 (3.9).

  4. gives ; gives . With , and , . At : , , a local minimum with . At : , , a local maximum with . At and : , saddles, with and .

  5. A is correct. By (7.6), the minimum variance rises by about . Since , , and exactly ; the difference is second order. B takes for ; C drops the 2 in .

  6. C is correct. The numerator of (7.7) for asset 1 is negative when . Only 0.60 exceeds the threshold.

  7. . The numerators of (7.10) are and , which sum to 0.000864. So . With and , (7.11) gives and , above both.

  8. . From (7.12), and , so . Then , as required. The model growth is .

  9. (a) and , so and : short the leveraged fund. , so . (b) Long-only, the constraint binds and the optimum holds only the index fund, . Here , so the multiplier is . (c) Allowing a 1% short in the leveraged fund () would cut the variance by about 0.00072 (exactly 0.000715). The long-only rule is what stops the investor from hedging the index fund with the leveraged one.

  10. Unconstrained, when , with total exposure 6.5: the cap binds. With only the cap binding, stationarity gives and , so , and . Then , both exposures are positive, and the sign constraints are slack. The excess growth is , against 20.50% unconstrained. At the margin the cap is worth 3.6 points per unit; raising it from 2 to 3 adds exactly 3.20 points.

  11. B is correct, by (7.9): . A is false, as the case shows. C is false. With 16% volatility and ten years of data, the mean annual return has a standard error of about points. Dividing by turns that into an error of about ±2 in a single asset's Kelly fraction.

  12. The previous year's weight beat the 50/50 mix in 6 of the 9 years, 2018–2026. Average volatility was 10.84% with the previous year's weight, 11.64% for 50/50 and 10.07% for each year's ex-post best weight, so estimation captured about half of the gain that perfect foresight would have delivered. It lost to 50/50 in 2019, 2023 and, by 0.002 points (14.253% against 14.251%), 2020.

Glossary

Term 中文 Meaning
Covariance 共變異數 Average product of two returns' deviations from their means.
Mixed partial derivative 混合偏導數 ; equal to under continuity.
Gradient 梯度 Vector of partial derivatives; direction of steepest ascent.
Hessian 海森矩陣 Matrix of second partial derivatives; curvature in every direction.
Positive definite 正定 for every .
Convex function 凸函數 Positive semidefinite Hessian everywhere; local minima are global.
Lagrangian 拉格朗日函數 ; its critical points give constrained optima.
Shadow price 影子價格 Rate at which the optimal value improves as a constraint is relaxed, (7.6).
Envelope theorem 包絡定理 An optimal value's derivative in a parameter equals the Lagrangian's partial derivative there.
KKT conditions KKT 條件 Stationarity, feasibility, non-negative multipliers, complementary slackness, (7.13).
Binding (active) constraint 緊限制式(有效限制) Inequality that holds with equality at the optimum.
Minimum-variance portfolio 最小變異數投資組合 Fully invested mix with the lowest variance, (7.7).
Tangency portfolio 切點投資組合 Fully invested mix with the highest Sharpe ratio, (7.10).
Kelly fractions 凱利比例 Growth-maximizing exposures, (7.12).
In-sample / out-of-sample 樣本內/樣本外 Data used to estimate a rule versus data used to test it.

References

  • Markowitz, H. (1952). "Portfolio Selection." Journal of Finance 7 (1): 77–91.
  • Merton, R. C. (1972). "An Analytic Derivation of the Efficient Portfolio Frontier." Journal of Financial and Quantitative Analysis 7 (4): 1851–1872.
  • Kelly, J. L., Jr. (1956). "A New Interpretation of Information Rate." Bell System Technical Journal 35 (4): 917–926.
  • Karush, W. (1939). "Minima of Functions of Several Variables with Inequalities as Side Conditions." M.Sc. thesis, Department of Mathematics, University of Chicago.
  • Kuhn, H. W., and A. W. Tucker (1951). "Nonlinear Programming." In J. Neyman (ed.), Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, 481–492. Berkeley: University of California Press.
  • Kjeldsen, T. H. (2000). "A Contextualized Historical Analysis of the Kuhn–Tucker Theorem in Nonlinear Programming: The Impact of World War II." Historia Mathematica 27: 331–361.
  • Boyd, S., and L. Vandenberghe (2004). Convex Optimization. Cambridge University Press, chapter 5 (§5.5 optimality conditions).
  • Nocedal, J., and S. J. Wright (2006). Numerical Optimization, 2nd ed. New York: Springer.
  • Kraft, D. (1988). "A Software Package for Sequential Quadratic Programming." Technical Report DFVLR-FB 88-28, DLR German Aerospace Center.
  • Chopra, V. K., and W. T. Ziemba (1993). "The Effect of Errors in Means, Variances, and Covariances on Optimal Portfolio Choice." Journal of Portfolio Management 19: 6–11.
  • Stewart, J., D. K. Clegg, and S. Watson (2021). Calculus: Early Transcendentals, 9th ed. Cengage, chapter 14 (partial derivatives, gradients, extreme values and Lagrange multipliers).
  • Yuanta Securities Investment Trust. 00679B fund information (tracking index ICE US Treasury 20+ Year Bond Index; listing 17 January 2017) and holdings report, retrieved 10 October 2026: yuantaetfs.com, 00679B.
  • 潘姿羽, 中央社 (30 December 2022). "新台幣全年重貶 9.83% 創 25 年最大跌幅,專家點出兩原因." TechNews 科技新報 (USD/TWD 30.708 at end-2022, up 3.018 on the year): finance.technews.tw.
  • SciPy 1.16.0 Release Notes, scipy.optimize (SLSQP ported to C; constraint multipliers returned): docs.scipy.org.
  • FinMind open data: TaiwanStockPrice, TaiwanStockDividendResult (0050, 00679B), retrieved 10 October 2026.