Optimization in One Variable
Find the top, prove it is the top, and measure how little its exact location matters.
| Mastery | After this module you should be able to: |
|---|---|
| 1identify the critical points of a differentiable function and classify them with the first- and second-derivative tests | |
| 2use concavity to show that a point where the derivative vanishes is the global maximum, and prove that the empirical log-growth function of a levered position is strictly concave | |
| 3solve maximization problems on a closed interval, recognising corner solutions created by leverage caps and long-only rules, and interpret the derivative at a corner as a shadow price | |
| 4derive the growth-optimal position of a binary bet, , and distinguish position size from risk size | |
| 5apply bisection and Newton's method to equations without closed-form solutions, such as a yield to maturity and the empirical Kelly condition, and compare their rates of convergence | |
| 6evaluate how sensitive an optimum and its value are to sizing errors and parameter changes, using the flat-top approximation and the envelope theorem | |
| 7estimate the in-sample growth-optimal leverage of a return series in Python, and explain why it is not a recommendation |
1Introduction
Optimization chooses the best value of a decision variable: how many lots to buy, which leverage to hold, what yield prices a bond. In one variable the tools are elementary: a zero derivative at the top, the second derivative to tell tops from bottoms, and the endpoints. A sound answer settles three questions: is the critical point the global top, does a constraint bind, and how much does the exact location matter? This module settles all three for the sizing problem every leveraged trader faces.
Take 0050's 2,902 daily total returns from 3 November 2014 to 8 October 2026. Suppose you had held a constant leverage : units of 0050 for each unit of equity, rebalanced daily and financed at zero cost. Which makes wealth grow fastest?
Wealth after days is , so the question is which maximizes the sample mean of . The answer is 5.42. At that leverage one unit of equity would have grown to about 1,122, against 10.0 for 0050 itself.
Before acting on it, read the rest of Exhibit 1. The same path fell 93.3% from peak to trough, and on 7 April 2025 alone it lost 54.2%. Half the leverage would have kept about three-quarters of the growth rate.
The textbook rule gives 5.68: close, but not equal. And the mean return is estimated with enough error that its 95% confidence band maps into anything from 2.7× to 7.6×.
Which of these numbers, if any, should guide a decision? By Section 7 you will derive each of them. You will see why the optimum is unique, and why its value is robust while its location is not. And you will see why the in-sample 5.42× describes the past rather than recommends a position.
| Leverage | Annual log growth | Growth of 1 unit | Annualized volatility | Maximum drawdown | Worst day (7 Apr 2025) |
|---|---|---|---|---|---|
| 1 (0050 itself) | 19.31% | 10.0 | 19.3% | −33.9% | −10.0% |
| 2 (exact daily 2×) | 34.87% | 64.2 | 38.6% | −57.9% | −20.0% |
| 2.71 (half Kelly) | 43.63% | 182.8 | 52.3% | −70.1% | −27.1% |
| 5.42 (full Kelly) | 58.84% | 1,122.5 | 104.7% | −93.3% | −54.2% |
| 9.74 (zero growth) | 0.00% | 1.0 | 188.0% | −99.99% | −97.4% |
TaiwanStockPrice and TaiwanStockDividendResult, retrieved 10 October 2026; total returns with reinvested dividends, adjusted for the 4-for-1 split of 18 June 2025, on the trading days common to 0050 and 00631L as in LM3. Computations: data/lm04_stats.py and the companion notebook LM04_lab.ipynb.Sections 2–4 build the tools: first- and second-order conditions, concavity and global optima, and constraints that create corner solutions. Section 5 derives the binary Kelly bet exactly. Section 6 solves equations with no closed form — the empirical Kelly condition and a bond's yield — by bisection and Newton's method. Section 7 measures the sensitivity of an optimum, introduces the envelope theorem and resolves the case; Section 8 does it all in Python.
is the leverage, units of the risky asset per unit of equity: a return becomes on equity, with free financing unless stated. is the expected log growth per period, its sample average, the maximizer, , and the fraction of full Kelly. For a sample of returns, and are the sample mean and standard deviation ( divides by ). , and are the means of , and , so in a model and in a sample.
Generic functions are and of a real variable , generic intervals and domains . A star marks a solution: is a critical point, a maximizer or a root. A binary bet (Section 5) wins with probability , , and gains or loses per unit staked.
Counts: returns; trades with wins (Section 5). Indices: for days or payment dates, for the steps of an iterative method, for the order of a derivative.
2First- and Second-Order Conditions
Optimization starts with a function of one decision variable, such as growth against leverage, and asks where it is largest. Calculus answers in two steps. The first-order condition locates the candidates; the second-order condition sorts them.
Local and global extrema
A point of the domain of a function is a global maximum of on if for every . It is a local maximum if for every in some open interval around . A maximum is strict if the inequality is strict for . Minima reverse the inequalities; maxima and minima together are extrema.
Every global maximum is a local one, but not conversely. A trader cares about the global maximum, while calculus sees only local information: the slope and curvature at a point. Section 3 shows when local information is enough.
The first-order condition
If has a local extremum at an interior point of its domain and is differentiable at , then .
Proof. Let be a local maximum. For small , , so the right difference quotient is and its limit gives . For small the numerator is still but the denominator is negative, so the left quotient is and . Hence ; a minimum reverses the signs.
An interior point of the domain at which , or at which does not exist, is a critical point of .
Three warnings follow from the proof. The condition is necessary, not sufficient: has and no extremum. It applies only at interior points; at an endpoint the one-sided argument yields an inequality instead (Section 4). And a kink, such as at 0, is a candidate even though no derivative exists there.
The second-order condition
Let and let be continuous near . If , then is a strict local maximum; if , a strict local minimum. If , the test is inconclusive.
Proof. Taylor's theorem with the Lagrange remainder (LM3, eq. (3.5)) gives for some between and . The linear term vanishes by the first-order condition. If , continuity keeps on a small interval around , so for every small . The case is symmetric.
When , look at higher derivatives. If the first non-zero derivative at has order , a Taylor expansion to order fixes the sign of . For even the point is an extremum, a maximum if ; for odd the graph crosses its tangent, with no extremum. Exhibit 2 collects the cases.
| Condition at | Conclusion | Example at |
|---|---|---|
| , | strict local maximum | |
| , | strict local minimum | |
| , at | no extremum (inflection point) | |
| , at | strict local maximum | |
| does not exist | compare the signs of on each side | (maximum) |
Growth as a function of leverage
Now the decision this module is about. A position of units of an asset per unit of equity turns a return into ; one period multiplies wealth by . Wealth compounds, so the natural objective is the expected log growth per period,
Section 5 shows why maximizing maximizes long-run wealth; here only its shape matters. Expand (LM3, eq. (3.3)) with , take expectations, and use (the population form of LM3's identity (3.7)):
The neglected terms are of third order in . This quadratic model is a downward parabola. Its first-order condition gives
and confirms a strict maximum. The second step drops , which is smaller than by the factor ; for 0050's daily returns . The Sharpe ratio is the mean return in excess of a risk-free rate , per unit of standard deviation (LM38 studies its estimation). With , as here, it is per period.
At the optimum : maximal growth is about half the squared Sharpe ratio per period. There the drag term is exactly half the levered mean , so at the optimum, volatility drag consumes half the levered arithmetic return. On 0050 at the exact optimum, a levered arithmetic mean of 114.83% a year becomes log growth of 58.84%: 51% is kept. LM42 derives the continuous-time versions, and .
(1) Over 2014–2026, 0050's mean daily total return was , with population standard deviation . Apply (4.3), and explain why it overshoots the exact optimum of Section 6. (2) A trading rule gains 2% or loses 1% of the position with equal probability. Apply (4.3) and compare with the exact optimum of Section 5, 25×.
(1) The mean square is , so ; dropping gives . The predicted maximal growth is a day, or 59.85% a year with . The exact in-sample optimum (Section 6) is 5.42, with 58.84% a year. The rule's leverage is 4% too high, while the growth it predicts is only 1.0 percentage point too high.
The overshoot comes from fat tails. Two more terms of add to the first-order condition. For normal returns and , up to terms smaller by the factor , so at the two terms cancel.
0050's daily returns have an excess kurtosis of 8.12 (LM19 develops this measure of tail weight). The mean of , divided by , exceeds the normal value 3 by that much. 0050's is 3.7 times that of normal returns with the same mean and variance, so the quartic term wins and pulls the optimum down. The two ±10% days are not the main cause: without them the rule still overshoots, 5.92 against an exact 5.74, three-quarters of the full-sample gap.
(2) Here and , so with predicted growth per trade, against the exact 25× and 5.889%. The rule fails because at the levered outcomes are +40% and −20%. That is far outside the range where is accurate (see the pitfall on large returns in LM3).
Setting finds every critical point where is differentiable: maxima, minima and inflection points alike. The tests of Exhibit 2, or concavity on the whole domain (Section 3), tell them apart. Check one of them before calling a critical point optimal.
可微函數在區間內部的極值點,切線一定是水平的(一階條件 ),但切線水平不保證是山頂:可能是谷底,也可能像 在 0 那樣只是反曲點。二階條件看彎曲方向, 向下彎就是局部最大。凱利的二次近似 是開口向下的拋物線,所以一階條件給出的 就是它的最大值。但它只在 不大時準確:+2%/−1% 的賭局開 20 倍,結果變成 +40%/−20%,二次近似就偏離真正的 25 倍。0050 的二次近似高估約 4%,主因是日報酬的厚尾(超額峰態 8.12);去掉那兩天 ±10%,高估仍剩約四分之三。
Find and classify the critical points of .
Answer
, so the critical points are and . With : , a strict local maximum with ; , a strict local minimum with . Neither is global on , because as .
3Concavity and Global Optima
The second-derivative test is local: it compares only with its neighbours. A sizing decision needs a global answer — no leverage anywhere does better. Concavity converts local information into that global guarantee.
A function is concave on an interval if, for all and ,
For twice-differentiable functions there is a working test, stated here without proof (Simon and Blume 1994, chapter 21). On an open interval, is concave exactly when throughout, and implies strict concavity. The converse of the last statement fails: is strictly concave although its second derivative vanishes at 0.
Let be differentiable and concave on an interval , and let be an interior point of with . Then is a global maximum of on . If is strictly concave, is the only maximizer.
Proof (when on ). For any , Taylor's theorem (3.5) gives , because and . If on , the inequality is strict for . The Deep Dive proves the theorem without second derivatives.
One corollary is used constantly. A differentiable, strictly concave function has at most one critical point, and if it has one, that point is the global maximizer. Finding the optimum then becomes a pure root-finding problem, the subject of Section 6.
Deep DiveA concave function lies below its tangent linesoptional · click to expand
Let be concave and differentiable on , with . For , concavity with weight on gives . Subtract and divide by :
As the left side tends to . Hence for every : the graph lies below each of its tangent lines. If , this reads , which proves the theorem with no assumption on .
The log-growth function is strictly concave
Apply this to the sample version of (4.1). For returns , the empirical growth function is
defined on the set of leverages with for every . If the sample contains both gains and losses, is the open interval . Differentiating term by term,
Suppose a sample contains at least one gain and one loss. Then is strictly concave on and has exactly one maximizer , the root of .
Uniqueness follows from (4.6), since is continuous and strictly decreasing on . For existence, let rise towards : the worst day's term tends to , so . As falls towards , the best day's term tends to , and so does .
The intermediate value theorem says that a continuous function that is positive at one point and negative at another is zero somewhere between them. It is stated here without proof (Stewart, Clegg and Watson 2021, chapter 2). So crosses zero, and exactly once. The same edges show that : near the worst day all but wipes out the account.
For 0050 the largest daily loss and gain were both exactly 10%: −10.00% on 7 April 2025 and +10.00% on 31 July 2026. Ten per cent is Taiwan's daily price limit, in force since 1 June 2015 (it was 7% before). So , and at 10× a single limit-down day ends the experiment. Exhibit 3 plots for 0050 against the quadratic model.
A familiar rule says growth returns to zero at about twice the Kelly leverage. Check it on 0050, and explain the gap using the worst day in the sample.
The quadratic model's zero is , outside . The exact zero-growth leverage, the root of above , is 9.74, or 1.80 times . It has no closed form; Section 6 shows how to find it.
The largest daily losses explain the gap. At , 7 April 2025 alone multiplies equity by , a log return of −3.64. That single day cancels the gains of the other 2,901 days.
At the same day costs −0.78 log points out of a total of 7.02. The logarithm punishes a large loss far more than it rewards an equal gain: , but . Fat tails therefore pull the growth curve down faster than any parabola as leverage approaches the ruin point.
"Local equals global" is a property of the objective, not of the search. A backtest's Sharpe ratio is usually bumpy in its parameters; a grid search returns the highest bump, which is largely noise (LM39). Kelly sizing is well posed because the logarithm of an affine function of is concave. Treat the optimum of a bumpy objective as a hypothesis.
凹函數(concave,圖形向下彎、弦在圖形下方)有個關鍵性質:可微的凹函數只要找到切線水平的點,它就是全域最大;嚴格凹時還是唯一的。平均對數成長 是許多 的平均,整體嚴格凹,所以只要樣本有漲有跌,凱利最適槓桿一定存在,而且只有一個。0050 單日最大漲跌都是 10%(台股漲跌幅限制),槓桿只能落在 −10 到 10 倍之間;越靠近 10 倍,一根跌停就讓帳戶歸零,曲線掉得比拋物線快得多。用語提醒:本書 concave=凹函數、convex=凸函數;有些微積分課本說的「凹口向上」其實是 convex。
Show that the growth function of a binary bet, with , and , is strictly concave on .
Answer
and . Both terms are negative on the interval, so and is strictly concave. Any critical point is therefore the unique global maximum, which Section 5 computes.
4Constraints and Corner Solutions
Real positions face limits. A cash account cannot borrow, many mandates cannot sell short, and brokers cap margin. Each limit confines to an interval, and the optimum may then sit on the boundary, where the derivative need not be zero.
A continuous function on a closed, bounded interval attains a maximum and a minimum on it.
The theorem is stated without proof (Stewart, Clegg and Watson 2021, chapter 4). It turns optimization on into a finite search, the closed-interval method. Evaluate at its critical points in and at both endpoints, and take the largest value; the method needs continuity only, not concavity.
First-order conditions at a corner
At an endpoint the argument of Fermat's theorem runs on one side only, so it delivers an inequality. If a differentiable attains its maximum on at , every left difference quotient is , so . Likewise a maximum at requires . Collecting the three cases,
These are the one-variable Karush–Kuhn–Tucker (KKT) conditions, which LM7 extends with Lagrange multipliers. For a concave they are also sufficient, by the tangent-line inequality of the Deep Dive in Section 3.
The derivative at a binding corner has an economic reading. If the optimum is and , relaxing the limit to raises the objective by about when is small. That derivative is the shadow price of the constraint: the rate at which relaxing it raises the objective, a marginal value rather than a total. In LM7 it becomes the Lagrange multiplier.
Let be differentiable and concave on an open interval containing , with an unconstrained maximizer . Then the clipped point maximizes on ; if is strictly concave, it is the only maximizer.
Proof. If there is nothing to show. If , then is non-increasing with , so on and is non-decreasing there. Hence is a maximizer, and the only one if is strictly concave, since then on . The case is symmetric.
Two corners in the 0050 data
A long-only investor without a margin account can still reach any leverage between 0 and 2. Mixing cash, 0050 and the daily 2× fund 00631L (LM3) does it, if rebalanced daily and ignoring costs and replication basis. The feasible set is . Exhibit 4 shows what in-sample optimization recommends year by year, with and without that interval.
| Year | 0050 total return | (bp a day) | (% a day) | Unconstrained | Optimum on |
|---|---|---|---|---|---|
| 2015 | −6.31% | −2.13 | 1.044 | −1.93 | 0 |
| 2016 | +19.64% | +7.81 | 0.956 | +8.37 | 2 |
| 2017 | +18.13% | +6.95 | 0.592 | +19.41 | 2 |
| 2018 | −4.95% | −1.50 | 1.046 | −1.40 | 0 |
| 2019 | +33.52% | +12.26 | 0.782 | +19.73 | 2 |
| 2020 | +31.08% | +12.12 | 1.462 | +5.44 | 2 |
| 2021 | +21.97% | +8.77 | 1.126 | +6.97 | 2 |
| 2022 | −21.34% | −8.79 | 1.389 | −4.46 | 0 |
| 2023 | +27.40% | +10.53 | 0.894 | +13.58 | 2 |
| 2024 | +48.67% | +17.66 | 1.580 | +5.95 | 2 |
| 2025 | +36.85% | +14.43 | 1.575 | +5.11 | 2 |
| Nov 2014 – Oct 2026 | +902.25% | +8.71 | 1.239 | +5.42 | 2 |
Two facts stand out. Year by year the unconstrained optimum ranges from a 4.46× short position to 19.73× long. And every constrained optimum is a corner: 0 in the three years with , 2 in the eight with .
(1) For the full sample, find the optimal leverage on , the growth it delivers and the shadow price of the ceiling. Estimate the gain from raising the ceiling to 2.5. (2) For 2022 alone, show that the long-only floor binds.
(1) is concave and , so by clipping the optimum is the corner : everything in the 2× fund. Growth is 34.87% a year (×64.2 over the sample), against 58.84% at . The corner condition holds: a year per unit of leverage, a positive number. That is the shadow price.
Raising the ceiling to 2.5 should add about a year; the exact gain is 6.37%, smaller because falls as rises. In practice 00631L grew ×44.5, not ×64.2, because of the costs and replication basis analysed in LM3.
(2) In 2022, bp a day and the unconstrained optimum is , a 4.46× short position. On the optimum is the corner , and the corner condition holds. In hindsight, cash beat every long position that year.
At a binding constraint the derivative is not zero, and solving returns an infeasible point: 5.42 when the ceiling is 2. Solve the unconstrained problem, then clip (for a concave objective) or compare endpoints (in general). Exhibit 4 also shows a benefit of corners: once a ceiling binds, the decision is the same for every estimate of above it. Corners make decisions robust to estimation error, a point Section 7 returns to.
現實中的槓桿有上下限:現金帳戶只用 0050 加 00631L,最多約 2 倍;不能放空,下限就是 0。閉區間上的最大值可能落在端點(角解),這時導數不一定為零。在上限 2 倍處 ,代表「如果能再多借一點,成長率還會增加」;這個導數就是限制的影子價格,是放寬一點點時的邊際增加速度,不是總增加量,也是 LM7 拉格朗日乘數的一維版本。目標函數是凹函數時,把無限制的最適解截到區間內就是答案。2015–2025 每個完整年度的答案都是角解:平均日報酬為負的三年落在 0,其餘八年的 都超過 2,落在上限 2。
Find the global maximum and minimum of on .
Answer
at , and the candidates give , , and . The global maximum is the corner , even though is a strict local maximum. The global minimum, −2, is attained twice: at the corner and at the interior point . The function is not concave, so clipping does not apply and every candidate must be compared.
5The Binary Kelly Bet
The simplest sizing problem has one source of risk, two outcomes and an exact solution. It exposes the confusion between position size and risk size, and shows how sharply the optimum depends on the win probability. LM41 extends it to general discrete bets.
Why maximize expected log growth
A trade gains a fraction of the position with probability and loses a fraction with probability , independently of other trades. With a position of times equity, one trade multiplies wealth by or . After trades with wins,
By the law of large numbers (LM17), , so the per-trade log growth converges to
If , the ratio of the two wealth paths behaves like and grows without bound. The fraction with the highest therefore ends up with more wealth than any other fixed fraction, with probability one (Breiman 1961; Thorp 2006). That is Kelly's (1956) criterion: maximize the expected log growth. MacLean, Thorp and Ziemba (2011) collect the classic papers on it.
Deriving the Kelly position
From (4.7),
The growth function is strictly concave (the Knowledge Check of Section 3), so the root of is the unique global maximum. Setting and cross-multiplying gives , so , and
No approximation was made: (4.8) is exact. Substituting it back gives and , so the maximal growth rate is
A rewriting makes (4.8) easier to read. Let be the break-even probability, at which the expected return is zero. Then
The Kelly position is proportional to the probability edge , with slope . It is positive exactly when the expected return is positive. With no edge, is an interior optimum with ; with a negative edge, a long-only rule gives the corner solution .
For a bet that gains with probability and loses with probability , the growth-maximizing position is exactly times equity, where . The growth function is strictly concave, so is unique, and the maximal growth (4.9) is positive whenever .
Position size versus risk size
On a losing trade the account loses of equity, the risk fraction. Multiplying (4.8) by ,
When the whole stake can be lost, and position equals risk. Then (4.11) is the classic Kelly formula for a bet paying to 1, which Thorp (2006) writes as . A stop-loss trade loses only of the position, so the two numbers differ by the factor : 100 for a 1% stop.
Equation (4.11) carries a second message. Scaling and by the same factor leaves the payoff ratio , , and unchanged, so by (4.9) it leaves unchanged. Only the payoff ratio and the win probability determine the attainable growth. The stop distance sets the position size that delivers it.
A trade gains 2% or loses 1% of the position with equal probability. Find the Kelly position, the risk per trade and the maximal growth. Then evaluate a trader who reads "Kelly = 25%" as "put 25% of capital into the trade".
By (4.8), : a position of 25 times equity. A loss costs of equity, the risk fraction of (4.11): . Growth at the optimum is per trade, from a levered arithmetic mean of .
The trader who holds 0.25 times equity risks only 0.25% per trade. Growth is per trade, 2.1% of the maximum. "25× equity" and "25% at risk" are the same decision in two units; mixing the units costs a factor of 100.
The same curve has two more landmarks. Its domain is , so one loss at 100× ends the account. And growth returns to zero at exactly , because .
Sensitivity to the win probability
Equation (4.10) makes the sensitivity explicit: , which is 150 for the +2%/−1% bet. One percentage point of win probability moves the Kelly position by 1.5 times equity. Exhibit 5 follows the bet as its true win probability falls, and asks what happens to a trader who keeps the 25× position.
| Win probability | Kelly position | Risk | Maximal growth per trade | Growth at a fixed 25× |
|---|---|---|---|---|
| 0.60 | 40.0 | 40.0% | 14.834% | +12.821% |
| 0.55 | 32.5 | 32.5% | 9.856% | +9.355% |
| 0.50 | 25.0 | 25.0% | 5.889% | +5.889% |
| 0.45 | 17.5 | 17.5% | 2.924% | +2.423% |
| 0.41504 | 12.3 | 12.3% | 1.451% | 0.000% |
| 0.40 | 10.0 | 10.0% | 0.971% | −1.042% |
| 1/3 (break-even) | 0 | 0 | 0 | −5.663% |
Growth at a fixed 25× stays positive until falls to 0.415. There the true Kelly position is 12.26, so the trader holds about twice Kelly, the zero-growth point of Section 7. Below it the trader loses money in the long run even though every trade still has a positive expected return. At the expected return is per unit of position, yet the 25× position shrinks wealth by 1.04% per trade.
Deep DiveKelly's information interpretation: maximal growth is a relative entropyoptional · click to expand
Write (4.9) with and :
Allowing (betting against the trade), maximal growth is positive whenever the true probability differs from the break-even one. It measures that difference in information units. This is the "new interpretation of information rate" of Kelly (1956); Cover and Thomas (2006, chapter 6) develop it for horse races.
The same decision can be stated as a 25× position or as 25% of equity at risk. Read as "25% of capital in the trade", it becomes a position one hundred times too small. Read as "25× leverage" on a trade with a 10% stop, it puts 250% of equity at risk: one loss is ruin.
State a Kelly number with its unit: a multiple of equity, or the fraction of equity lost if the trade fails. Recompute the position whenever the stop distance changes.
凱利公式 算的是「部位是權益的幾倍」,風險公式 算的是「一筆交易賠掉時損失權益的比例」(整筆可能賠光、 時就是經典的 ),兩者差一個 。+2%/−1%、勝率五成:部位 25 倍、每筆風險 25%,說的是同一件事;把「25%」誤讀成「拿 25% 資金下單」,部位只剩凱利的百分之一,成長率只剩最大值的約 2%。另一個重點是對勝率極度敏感:,勝率每掉 1 個百分點,部位就少 1.5 倍權益。固定開 25 倍而勝率掉到 0.415 時,每筆交易的期望報酬仍是正的,長期複利成長率卻已歸零。
A trade gains 3% with probability 0.45 and otherwise loses 1.5%. Find the Kelly position and the risk per losing trade, and verify the risk with (4.11).
Answer
times equity. A loss costs of equity. By (4.11), : the same 17.5%.
6Root Finding: Bisection and Newton's Method
The binary bet is the exception: most first-order conditions have no closed form. The empirical Kelly condition is the sample version of . Once denominators are cleared, for 0050 it becomes a polynomial equation of degree 2,709, one less than the number of distinct non-zero daily returns.
A bond's yield and a project's internal rate of return are roots of polynomials too. Numerical root finding solves to any accuracy, and for a smooth concave objective , maximizing is root finding on .
Bisection
Let be continuous on , with and of opposite signs. Evaluate at the midpoint and keep the half-interval whose endpoints still have opposite signs. Repeat.
By the intermediate value theorem (Section 3) every interval in the sequence contains a root. After steps the bracket has width , so its midpoint is within of a root. A bracket of width takes
steps: one binary digit per step, or about 3.3 steps per decimal digit. The error bound shrinks by a constant factor each step, which is called linear convergence. Bisection needs only continuity and a sign change, and it cannot fail.
Newton's method
Newton's method replaces by its tangent at the current point, the first-order Taylor polynomial of LM3, and jumps to the tangent's root:
Applied to for a maximization, the step is : each iteration maximizes the second-order Taylor model of at .
The speed comes from the remainder. Let be a root with and write . Taylor's theorem gives for some between and . Dividing by and using (4.12),
The error is squared at each step: quadratic convergence. Once is small, the number of correct digits roughly doubles per iteration. The constant says how small is small: squaring helps once .
Let be twice continuously differentiable near a root with . Then there is an interval around such that Newton's method started anywhere in it converges to . Moreover, for some constant .
The theorem is stated without its full proof; (4.13) is the core, and Burden, Faires and Burden (2016, chapter 2) supply the details. It is a local result. Far from the root the tangent can point anywhere. Where the step is huge, and a step can leave the domain or land near a different root.
Safeguarded methods combine the two ideas. They keep a bracket, accept a fast step (Newton or secant) when it lands inside it, and bisect otherwise. Brent's (1973) method, available as scipy.optimize.brentq, is the standard version: it never loses the bracket and converges superlinearly near the root.
The empirical Kelly condition
Newton's method on uses (4.5) and (4.6). Start from , where and :
The quadratic rule (4.3) is exactly one Newton step from zero. Each later step refits the quadratic model at the current point. Exhibit 6 runs both methods on 0050. Bisection starts from the bracket , since and : falls towards as approaches 10, where it is undefined.
| Step | Newton | Newton error | Bisection bracket width | Bisection midpoint error |
|---|---|---|---|---|
| 0 | 0 | 5.4 | 9.9 | 0.47 |
| 1 | 5.651642810 | 4.95 | 2.0 | |
| 2 | 5.423914969 | 2.475 | 0.77 | |
| 3 | 5.421488036 | 1.2375 | 0.15 | |
| 4 | 5.421487783 | 0.61875 | 0.16 | |
| 10 | ||||
| 20 | ||||
| 30 | ||||
| 44 |
A five-year bond pays an annual coupon of 1.5 per 100 face value and trades at 97.80 (illustrative numbers). Its yield solves , where
Let , with . At the bond would trade at par, so , and ; hence . Next, and give , and prices the bond to within . The yield is 1.9663%.
Bisection starts from and . It needs steps against Newton's three. Newton is faster here because is smooth and gently curved and the coupon rate is a good starting point. Bisection's advantage is that it needs none of these.
Newton's method trusts the tangent. Cash flows of −100, +230 and −132 (a project with a closing cost) have two internal rates of return, 10% and 20%. Their NPV is flat at 14.8%. Newton started at 14% first jumps to −1.20% and ends at 10%; started at 15%, it jumps to 72.50% and ends at 20%.
On 0050's Kelly condition Newton converged from all 199 starting points from −9.9 to 9.9 (step 0.1), in at most eight steps. That is a property of the function, not of the method: use a bracketing method (brentq) whenever a sign change is known.
二分法每一步把區間砍半,穩但慢:每多一位小數大約要 3.3 步,只要函數連續、端點異號就一定收斂。牛頓法用切線逼近,誤差每一步大約平方一次,正確位數接近翻倍;但它需要好的起點,切線太平時會跳得很遠,可能跳出定義域或跑到另一個根。實務上兩者混合(scipy 的 brentq):用區間保證不跑掉,再用快速的步驟收斂。最妙的是:從 出發的第一步牛頓法,剛好就是二次近似 。在 0050 上牛頓法四步就到機器精度,二分法則要 44 步才把區間縮到 以下。
A yield is known to lie between 0% and 10%. How many bisection steps reduce the bracket to a width of (0.01 bp)?
Answer
The width after steps is , so we need , i.e. . Seventeen steps suffice.
7Sensitivity of the Optimum: Flat Tops and the Envelope Theorem
Two questions remain once an optimum is found: what does missing it cost, and how do the optimum and its value move with the inputs? The second-order term of a Taylor expansion answers the first, the envelope theorem the second. Together they resolve the case.
Flat tops
At an interior maximum the first-order term of the Taylor expansion vanishes, so
where the approximation drops terms of third order in . The loss from mis-sizing is of second order in the sizing error. An error twice as large costs four times as much, and a small error costs very little. The top of a smooth maximum is flat.
In the quadratic model the flat top has an exact form. Write (4.2) as , with and . Substituting ,
A position of times the Kelly leverage keeps the fraction of the maximal growth rate, with times the volatility. Half Kelly keeps 75% of the growth at half the volatility; quarter Kelly keeps 43.75% at a quarter. Over-betting mirrors under-betting: also keeps 75%, at three times the volatility of half Kelly. Growth is zero at and negative beyond.
This is Thorp's (2006) . Exhibit 7 compares (4.15) with the exact curves for 0050 and for the +2%/−1% bet.
Below full Kelly the three curves nearly coincide: half Kelly keeps 75% in the model, 74.2% for 0050 and 76.1% for the bet. Beyond full Kelly they part: the bet's curve is exactly symmetric about and reaches zero at , as Example 4 found. 0050 falls away: 1.5 times Kelly keeps 69.1% instead of 75%, 1.75 times keeps 17.9%, and growth is zero at 1.80 times. Over-betting a fat-tailed asset is far worse than the parabola suggests; under-betting is not.
Compare full and half Kelly on 0050 in-sample, and check the flat-top approximation (4.14) for the growth lost.
From Exhibit 1, half Kelly () earned 43.63% a year against 58.84%, keeping 74.2% of the maximal growth rate; the quadratic model says 75%. In exchange it halved the volatility (52.3% against 104.7%), halved the worst day (−27.1% against −54.2%) and cut the maximum drawdown from −93.3% to −70.1%. Terminal wealth was 183 against 1,122: a smaller multiple from a far more survivable path.
For (4.14), a day, or −0.0439 a year. The predicted loss at is a year, against an actual loss of . The quadratic term gets the size of the loss right to within one percentage point, 2.7 units of leverage away from the top.
The envelope theorem
Let be differentiable, and for each let be an interior maximizer of that depends differentiably on . Then the value function satisfies
Proof. By the chain rule, . The first-order condition sets at the maximizer, which leaves .
The theorem is the flat top seen from the side. When a parameter moves the optimizer moves too, but to first order this does not change the value. To price a parameter change, differentiate the objective with respect to the parameter at the old optimum; no re-optimization is needed. Simon and Blume (1994, chapter 19) give the multivariable version; two applications follow.
- The value of edge. In the quadratic model without the term, and , so . The envelope theorem gets there in one line: , evaluated at . The marginal value of edge is the position you hold.
- The cost of financing, worked out in Example 7.
Suppose the investor borrows (and lends idle cash) at a simple rate per trading day, so that one day multiplies wealth by . Use the envelope theorem to find how maximal in-sample growth responds to near zero. Check the answer for an illustrative financing rate of 2% a year, charged as per trading day.
With , the partial derivative is . At and the first-order condition does the work:
At 2% a year the first-order estimate is . Re-optimizing exactly gives and growth of 50.46% a year, a fall of 8.38%. Keeping the old leverage of 5.42 still earns 49.99%: re-optimizing is worth only 0.47% a year, which is precisely the envelope theorem's point.
How well the data locate the optimum
The flat top has a second face. If a wide range of leverages earns nearly the same growth, the data can barely tell those leverages apart. The value of the optimum is robust; its location is fragile.
For independent returns, the standard error of a sample mean is (stated without proof; LM20 derives it and develops confidence intervals). A 95% confidence interval for the true mean is standard errors, where 1.96 is the value a standard normal variable exceeds with probability 2.5%.
Turn the sampling error of into a 95% band for 0050's Kelly leverage. Shift all daily returns equally, putting the sample mean at each end of the 95% interval, and re-solve the exact Kelly condition. What would holding 5.42 achieve at either end of the band?
The standard error is a day against , a relative error of 26%. The 95% interval for the mean runs from 0.042% to 0.132% a day. On the shifted samples the exact Kelly leverage runs from 2.7 to 7.6. The quadratic rule would give 2.7 to 8.6, missing the fat tails at the high end.
The cost of the error is asymmetric. If the true Kelly leverage were 2.7, holding 5.42 would be , where (4.15) leaves almost nothing. On the shifted sample's exact curve, 5.42 even loses 0.63% a year. If the true value were 7.6, holding 5.42 would be , which keeps 92% by (4.15) and 90% on the shifted sample.
The data's own history agrees. Over trailing three-year windows (Exhibit 8) the in-sample optimum wandered from 0.94 to 8.87, a range of more than nine to one. The low came in the window ending 17 January 2024, the high in the window ending 13 August 2018. The latest window says 6.63.
Under-betting costs a little; over-betting can cost everything. That asymmetry, not timidity, is the case for fractional Kelly, a quarter to a half of . LM24 and LM44 formalize it, and LM45 adds parameter uncertainty. MacLean, Thorp and Ziemba (2010) summarize the trade-off: maximal long-run growth, but large positions and a real chance of losing most of the wealth.
In the quadratic model and , with SR and σ annualized (LM42). A backtest Sharpe ratio of 3 at 15% volatility suggests and log growth of 4.5 a year. That is a 90-fold gain every year, which should itself raise suspicion.
If the true Sharpe ratio is half the backtest's, 20× is twice Kelly and growth is zero. If it is 1, 20× is three times Kelly, and by (4.15) growth is a year. Wealth then falls 78% a year despite a genuine edge. Backtests overstate Sharpe ratios (LM38, LM39): size from a haircut estimate, then take a fraction of Kelly.
Resolving the case
- 5.42 is the unique global maximizer of in-sample growth, because is strictly concave (Section 3); Newton's method finds it in four steps from zero (Section 6).
- 5.68 is the quadratic rule , and 5.65, one Newton step from zero, its version with kept. Both overshoot because 0050's daily returns are fat-tailed (Section 2).
- 74.2% is the flat top: half Kelly kept almost three-quarters of the growth rate at half the volatility.
- 9.74, the zero-growth leverage, is 1.80 times rather than 2, pulled in by the largest losses (Section 3).
- 2.7 to 7.6 is the range of Kelly leverages consistent with the sampling error in ; the rolling windows wandered from 0.94 to 8.87.
The in-sample 5.42× therefore describes 2014–2026; it is not a recommendation. It treats the sample mean as truth, assumes free daily rebalancing and ignores financing, which costs 8.38% a year of growth at 2%. Even with perfect hindsight, it came with a 93.3% drawdown. LM20 and LM45 supply the statistics that must come before any leverage decision.
最大值附近曲線很平(一階項為零,只剩二階項):槓桿抓錯一點,成長率幾乎不變。在二次模型裡, 倍凱利保留 的成長率:半凱利保留 75%、波動減半,0050 實際是 74.2%。但「平」的另一面是資料很難分辨最適點在哪:0050 平均報酬的估計誤差約 26%,凱利槓桿的 95% 區間約 2.7 到 7.6 倍,滾動三年的樣本內最適值在 0.94 到 8.87 倍之間跳動。而且押太多比押太少危險得多,0050 超過 1.8 倍凱利,長期成長率就轉負。包絡定理則說:參數變動對「最大成長率」的一階影響,只要對參數偏微分、在原最適點取值即可,不必重新最佳化;例如融資利率每高 1 個百分點,最大成長率約少 個百分點。
In the quadratic model, what share of maximal growth does 1.5 times Kelly keep, and how does its volatility compare with half Kelly's? What does 0050's exact in-sample curve say?
Answer
, the same as half Kelly, at three times half Kelly's volatility: identical growth for triple the risk. On 0050's exact curve 1.5 times Kelly keeps only 69.1%, because fat tails punish over-betting more than the parabola does.
8Optimization in Python
SciPy's optimize module contains every method of this module. Three habits keep results honest: pass the domain explicitly, give a bracket whenever one is known, and compare at least two methods.
import numpy as np
import polars as pl
from scipy.optimize import brentq, minimize_scalar, newton
# 0050 total returns on days common with 00631L, from the lab's data/ folder
common = pl.read_csv("data/00631L_daily_raw.csv").select("date")
div = pl.read_csv("data/0050_dividends.csv")
px = (pl.read_csv("data/0050_daily_raw.csv").join(common, on="date")
.join(div, on="date", how="left").sort("date"))
split = pl.when(pl.col("date") == pl.lit("2025-06-18")).then(4.0).otherwise(1.0)
P, D = pl.col("close"), pl.col("cash_dividend").fill_null(0.0)
R = (px.select((P + D) * split / P.shift(1) - 1)
.to_series().drop_nulls().to_numpy())
def gbar(f): # mean log growth, eq. (4.4)
return np.log1p(f * R).mean()
def dg(f): # Kelly condition, eq. (4.5)
return (R / (1 + f * R)).mean()
def d2g(f): # curvature, eq. (4.6)
return -((R / (1 + f * R)) ** 2).mean()
lo, hi = -1 / R.max(), 1 / -R.min() # domain: 1 + f R_t > 0 for all t
f_brent = brentq(dg, 0.0, 0.999 * hi, xtol=1e-14) # bracketed root
f_newton = newton(dg, x0=0.0, fprime=d2g, tol=1e-14) # Newton from f = 0
f_bound = minimize_scalar(lambda f: -gbar(f), bounds=(0.0, 0.999 * hi),
method="bounded").x
print(f"domain ({lo:.2f}, {hi:.2f}) brentq {f_brent:.6f} "
f"newton {f_newton:.6f} bounded {f_bound:.4f}")
# domain (-10.00, 10.00) brentq 5.421488 newton 5.421488 bounded 5.4215
minimize_scalar with method="bounded" maximizes directly by Brent's derivative-free search. It stops at its default tolerance of about , which is why it is printed with fewer digits. The lab applies gbar to the fractional-Kelly curves of Exhibit 7 and traces Newton's method and bisection for the yield of Example 5.
Called without bounds, minimize_scalar(lambda f: -gbar(f)) returns nan on the 0050 data, and SciPy reports that it found no valid bracket. Explain why, and give two fixes.
Answer
The unbounded Brent search expands its trial points beyond the domain . There on some day, np.log1p returns nan or , and the comparisons that drive the search break down. Either restrict the search, minimize_scalar(..., bounds=(0.0, 0.999 * hi), method="bounded"), or solve the first-order condition with brentq on a bracket inside the domain.
The Polars-first notebook LM04_lab.ipynb rebuilds every computational exhibit and example and solves the computational practice problems.
The notebook reads FinMind files that you download once with your own token. From the repository root, run uv run python data/fetch_finmind.py, then uv run python data/make_lab_data.py. FinMind's licence does not allow the book to redistribute the data. After that, the notebook runs offline.
Summary
- At an interior maximum of a differentiable function the derivative is zero (Fermat's theorem). A negative second derivative confirms a strict local maximum; a zero second derivative decides nothing; use the tests of Exhibit 2.
- For a differentiable concave function a zero of the derivative is a global maximum, and strict concavity makes it unique. The empirical growth function is strictly concave, so the empirical Kelly leverage exists and is unique whenever the sample contains gains and losses.
- On a closed interval, compare critical points with endpoints. At a binding corner the derivative is the constraint's shadow price, a marginal rate; for a concave objective the constrained optimum is the unconstrained one clipped to the interval.
- The binary Kelly bet has the exact solution . It is a position size; the risk size is . The position moves by per unit of win probability, and attainable growth depends only on and the payoff ratio .
- Bisection gains a binary digit per step; Newton's method squares the error near a simple root but needs a good start. On 0050's Kelly condition, Newton's first step from zero is the quadratic rule, and four steps reach machine precision. Bisection needs 44 steps to shrink the bracket below .
- Near an optimum, growth is flat: keeps of the maximum in the quadratic model. The envelope theorem prices parameter changes without re-optimizing: at the empirical Kelly point, each unit of financing rate costs units of growth.
- In-sample, 0050's growth-optimal leverage over 2014–2026 is 5.42, with 58.84% a year of log growth and a 93.3% drawdown. Sampling error spans 2.7 to 7.6, and over-betting costs far more than under-betting: the number describes the past and is not a recommendation.
Practice Problems
-
The function on attains its maximum at:
- A.
- B.
- C.
-
Which statement about at is correct?
- A.The second-derivative test is inconclusive, and is a strict local minimum.
- B.The second-derivative test shows a local maximum.
- C. is an inflection point.
-
A strategy's annual log growth as a function of leverage is approximately . Find the Kelly leverage, the maximal growth, the growth at half Kelly and the leverage at which growth returns to zero.
-
A trade gains 5% with probability 0.55 and otherwise loses 5%. The Kelly position is closest to:
- A.0.10 times equity
- B.2 times equity
- C.11 times equity
-
Show that the Kelly position of a binary bet is positive if and only if the expected return per unit of position, , is positive. Show also that the maximal growth is then positive.
-
A trading rule wins +8% with probability 0.30 and stops out at −2%. (a) Find the Kelly position, the risk per trade and the maximal growth per trade. (b) If the true win rate is 0.25, by what factor is the position from (a) too large, and what growth per trade does it earn?
-
The number of bisection steps needed to shrink a bracket around a root to a width of is closest to:
- A.8
- B.19
- C.27
-
An investor's annual growth is approximately , but the account limits leverage to . Find the optimal leverage, the growth it delivers and the shadow price of the limit. What would raising the limit to 0.9 add?
-
In the quadratic model , use the envelope theorem to find , and verify the result by differentiating directly.
-
In the quadratic model, half Kelly keeps:
- A.75% of the maximal growth rate at 50% of the full-Kelly volatility
- B.50% of the maximal growth rate at 50% of the full-Kelly volatility
- C.75% of the maximal growth rate at 25% of the full-Kelly volatility
-
A backtest reports an annualized Sharpe ratio of 2.4 at 12% annualized volatility. (a) What leverage does full Kelly suggest in the quadratic model? (b) For which true Sharpe ratio does that leverage produce zero long-run growth? (c) What are the annual log growth and the annual wealth factor if the true Sharpe ratio is 0.8?
-
(Python) Using the lab data, compute the empirical Kelly leverage of 0050 from the returns of 2 January 2020 to 8 October 2026. Compare it with for the same window.
-
Expand the binary growth function (4.7) to second order around , and show that the resulting approximate Kelly position is . Evaluate it for the +2%/−1% bet at and explain why it falls short of the exact 25.
Solutions
-
C is correct. at , and gives . The maximum value is , which B mistakes for the maximizer. The endpoint , where , is the minimum on .
-
A is correct. and : the first non-zero derivative has even order and is positive, so is a strict local (indeed global) minimum (Exhibit 2). C would require the first non-zero derivative to have odd order, as for .
-
gives , and confirms the maximum. Maximal growth is , or 4.5% a year. Half Kelly gives , which is 75% of , as (4.15) predicts. Growth returns to zero at .
-
B is correct. . The risk per trade is , which is also with ; A mistakes that risk fraction for the position. C keeps only the first term, , and ignores the losses.
-
By (4.10), with , so if and only if , that is, , that is, . In that case is the unique maximizer of the strictly concave , and , so . (Alternatively, the Deep Dive shows for .)
-
(a) times equity, risking per trade; check: (4.11) gives . With and , per trade.
(b) At , , so 6.25× is exactly twice Kelly: by (4.10), halves with the edge over break-even (), from 0.10 to 0.05. Growth at 6.25× is per trade, 16.5% of the attainable 0.738%. Equation (4.15) predicts zero, but the exact zero is at 6.56×, 2.10 times Kelly. A rare large gain () makes growth fall more slowly beyond than it rises below it; at the curve is symmetric.
-
C is correct. , so 27 steps. A counts decimal digits instead of binary ones; B uses the natural logarithm, , instead of .
-
The unconstrained optimum is , so by clipping the optimum is the corner , with against an unconstrained 0.050. The shadow price is a year per unit of leverage: a marginal rate, not a total. Raising the limit to 0.9 adds about percentage points to first order and exactly points, because falls as rises. A limit above 1 adds nothing more: the whole gain on offer is 0.2 points.
-
; at this is . Directly, . The two agree, as (4.16) requires; the change in never enters.
-
A is correct. By (4.15) with , the growth kept is , and volatility scales with . B scales growth with the position, as if growth were linear in ; C confuses volatility with variance, which falls to 25%.
-
(a) . (b) Growth is zero at , so when the true Kelly leverage is 10, which means a true Sharpe ratio of : half the backtest's. (c) A true Sharpe ratio of 0.8 gives and , so growth is a year, a wealth factor of a year.
-
Over the 1,635 returns from 2 January 2020 to 8 October 2026, the empirical Kelly leverage is 5.38 and . As over the full sample (5.68 against 5.42), is about 5% too high. The window's optimum differs from the full-sample 5.42 by far less than the sampling error of Section 7.
-
Using with and : , a parabola with its maximum at . This is the quadratic rule (4.3) with and . For +2%/−1% at it gives , against the exact 25.
The neglected third-order term, , is positive here, so it raises growth at high leverage and pushes the true optimum up. The bet is symmetric about its mean (±1.5%), so with : positive because . A positive mean alone does not settle the sign. For +1% with probability 0.9 against −5%, but , and the quadratic rule overshoots: 11.8 against an exact 8.
Glossary
| Term | 中文 | Meaning |
|---|---|---|
| Critical point | 臨界點 | Interior point where or does not exist. |
| First-order condition | 一階條件 | at an interior extremum (Fermat's theorem). |
| Second-order condition | 二階條件 | Sign of at a critical point: negative for a strict local maximum. |
| Local / global maximum | 局部/全域極大值 | Best value in a neighbourhood / on the whole domain. |
| Inflection point | 反曲點 | Point where curvature changes sign; not an extremum. |
| Concave function | 凹函數 | Chords lie on or below the graph (); if is differentiable, a point with is a global maximum. |
| Convex function | 凸函數 | Chords lie on or above the graph (). |
| Intermediate value theorem | 中間值定理 | A continuous function that changes sign on an interval is zero somewhere in it. |
| Extreme value theorem | 極值定理 | A continuous function on attains its maximum and minimum. |
| Corner solution | 角解 | Optimum at a boundary of the feasible interval. |
| Shadow price | 影子價格 | Rate at which relaxing a binding constraint raises the objective. |
| KKT conditions | KKT 條件 | First-order conditions with inequality constraints; at a corner, a sign condition on . |
| Kelly criterion | 凱利準則 | Choose the position that maximizes expected log growth. |
| Kelly position | 凱利部位 | Growth-optimal position , as a multiple of equity. |
| Risk fraction | 風險比例 | Fraction of equity lost if a trade fails, . |
| Break-even probability | 損益兩平勝率 | : the win probability with zero expected return. |
| Fractional Kelly | 分數凱利 | A position with , such as half or quarter Kelly. |
| Bisection method | 二分法 | Halve a sign-change bracket at each step; linear convergence. |
| Newton's method | 牛頓法 | ; quadratic convergence near a simple root. |
| Quadratic convergence | 二次收斂 | : correct digits roughly double per step. |
| Envelope theorem | 包絡定理 | evaluated at the optimum. |
| Yield to maturity | 到期殖利率 | Discount rate that equates a bond's present value to its price. |
| Internal rate of return | 內部報酬率 | Discount rate at which net present value is zero. |
| Relative entropy | 相對熵 | ; equals the binary bet's . |
References
- Kelly, J. L., Jr. (1956). "A New Interpretation of Information Rate." Bell System Technical Journal 35 (4): 917–926.
- Breiman, L. (1961). "Optimal Gambling Systems for Favorable Games." In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Vol. 1, 65–78. University of California Press.
- Thorp, E. O. (2006). "The Kelly Criterion in Blackjack, Sports Betting, and the Stock Market." In S. A. Zenios and W. T. Ziemba (eds.), Handbook of Asset and Liability Management, Vol. 1, 385–428. Amsterdam: Elsevier.
- MacLean, L. C., E. O. Thorp, and W. T. Ziemba (eds.) (2011). The Kelly Capital Growth Investment Criterion: Theory and Practice. World Scientific Handbook in Financial Economics Series, Vol. 3. Singapore: World Scientific.
- MacLean, L. C., E. O. Thorp, and W. T. Ziemba (2010). "Long-Term Capital Growth: The Good and Bad Properties of the Kelly and Fractional Kelly Capital Growth Criteria." Quantitative Finance 10 (7): 681–687.
- Burden, R. L., J. D. Faires, and A. M. Burden (2016). Numerical Analysis, 10th ed., chapter 2, "Solutions of Equations in One Variable". Boston: Cengage Learning.
- Brent, R. P. (1973). Algorithms for Minimization without Derivatives. Englewood Cliffs, NJ: Prentice-Hall (Dover reprint, 2002).
- Stewart, J., D. K. Clegg, and S. Watson (2021). Calculus: Early Transcendentals, 9th ed., chapter 2 (continuity and the intermediate value theorem) and chapter 4, "Applications of Differentiation" (maximum and minimum values, optimization problems, Newton's method). Cengage (ISBN 9781337613927).
- Simon, C. P., and L. Blume (1994). Mathematics for Economists. New York: W. W. Norton (chapter 19 on envelope theorems; chapter 21 on concave functions).
- Cover, T. M., and J. A. Thomas (2006). Elements of Information Theory, 2nd ed., chapter 6, "Gambling and Data Compression". Hoboken, NJ: Wiley.
- 臺灣證券交易所. 《臺灣證券交易所60週年特刊》, 05 證券交易市場, p. 181: the daily price limit was widened from 7% to 10% on 1 June 2015. Online edition, page 181.
- FinMind open data:
TaiwanStockPriceandTaiwanStockDividendResult(0050, 00631L), retrieved 10 October 2026.