1Purpose and scope
OPM produces monthly voluntary prepayment projections (SMM and CPR — Conditional Prepayment Rate, the total prepayment speed; the voluntary component alone is the Conditional Repayment Rate, CRR) and credit projections (delinquency transitions and involuntary liquidation — the Conditional Default Rate, CDR) for U.S. residential mortgage collateral, at the single-loan, representative-line, and portfolio level, under user-specified interest-rate, home-price, and market scenarios. Seven products are covered, spanning both agency collateral (Fannie Mae and Freddie Mac) and government collateral (Ginnie Mae):
| Product | Collateral | Version | Estimation data |
|---|---|---|---|
| Conventional | Agency conforming fixed-rate, post-2012 (UMBS TBA) | v89.86_bal30frz | Fannie Mae single-family loan performance |
| FHA | Government fixed-rate, FHA-insured | v1.25 | Ginnie Mae monthly loan-level disclosure |
| VA | Government fixed-rate, VA-guaranteed | v2.19 | |
| USDA | Government fixed-rate, USDA/RHS-guaranteed | v1.10 | |
| ARM | Agency hybrid adjustable-rate (3/1 through 10/1) | v2.4 | Freddie Mac non-standard dataset |
| Government ARM | Ginnie Mae FHA and VA hybrid adjustable-rate | v1.0 | Ginnie Mae ARM disclosure panel |
| GNMA (G2 TBA) | Ginnie Mae II mixed FHA/VA/USDA collateral | composite | No additional parameters — see §7 |
The models are conditional: given a realized or assumed path of rates and home prices, they project prepayment and credit behavior. They do not forecast interest rates. Validation accordingly measures the accuracy of the behavioral response against realized history, not the foresight of any scenario.
Four deployment surfaces share a single computational core: an interactive single-loan forecaster (this site); a representative-line mode that expands disclosed pool characteristics into a weighted grid of constituent loans; a portfolio scenario engine; and an adapter implementing a third-party analytics platform's prepayment and loss function interface.
2Modeling approach ↑ top
Structural decomposition
Each fixed-rate model decomposes total voluntary prepayment (CRR) into four economically distinct components: housing turnover (present in all environments and only weakly rate-sensitive), rate refinance (driven by the borrower's rate incentive and gated by qualification), cash-out refinance (gated by accumulated equity, and structurally penalized when extracting equity requires resetting the full balance to an above-market rate), and curtailment (partial principal prepayments, which are income-driven and exhibit diminished response at both deep rate extremes). The reported decomposition reconciles exactly to the reported total in every month. This structure is the principal defense against regime confounding: an estimated effect must reside on the component whose economics produce it.
Constrained, interpretable parameterization
Components are constructed from parametric response curves on continuous characteristics (incentive, loan age, credit score, loan-to-value, balance), categorical multipliers, and seasoning profiles. Sign and range restrictions are embedded in the parameterization itself rather than imposed as penalties, so that economically inadmissible states — a negative maximum turnover rate, an inverted refinance response — are unreachable by the estimation procedure. The refinance component is specified as a mixture of two borrower populations with differing refinance thresholds; a loan that remains outstanding despite a persistent rate incentive is progressively re-classified toward the less responsive population. This makes refinance burnout an endogenous property of the specification and the incentive response structurally monotone. Of the conventional model's 250 registered parameters, 218 are estimated and 32 are structural values set by the modeler where the estimation objective cannot identify them; each of the latter is individually flagged in the parameter registry.
Forward-looking regime variables
Regime effects — the 2020–21 refinance wave, the 2022–25 turnover contraction — are driven exclusively by functions of the interest-rate path, such as the current rate relative to its trailing average and the duration of an elevated-rate state. Calendar and vintage indicator variables are excluded by design: an indicator fits history but conveys no information under a forward scenario, whereas a rate-path variable generalizes to any scenario by construction.
Program rules as dated policy, not estimated behavior
Government program rules — FHA mortgage insurance premium schedules, the VA seasoning, recoupment, and net-tangible-benefit requirements of 38 U.S.C. §3709, streamline refinance eligibility windows, funding-fee changes — enter the models as dated policy schedules with effective dates taken from the primary sources (Mortgagee Letters, All-Participants Memoranda, public laws), rather than as estimated coefficients on calendar time. Conversely, issuer and servicer operations (such as delinquent-loan buyout timing) are never estimated as borrower behavior; the credit model explicitly reconstructs borrower outcomes from behind buyout censoring (§8).
3Data ↑ top
Loan performance data
| Book | Source | Observation span | Scale |
|---|---|---|---|
| Conventional | Fannie Mae single-family loan performance files, 56 acquisition-vintage files, 2011Q4–2025Q3 | Jan 2012 – Sep 2025 | 1.67 billion loan-months; 16.5 million payoff events |
| FHA / VA / USDA | Ginnie Mae monthly loan-level disclosure (Ginnie Mae I and II) | Oct 2013 – Jul 2026 (154 consecutive disclosure months) | FHA segment 1.08 billion loan-months |
| ARM | Freddie Mac non-standard dataset, hybrid ARMs 3/1–10/1 (6.4 million origination records) | Mar 1999 – Sep 2021 (performance history ends at the vendor's cutoff) | 81.7 million loan-months |
| Government ARM | Ginnie Mae ARM disclosure; contract terms (caps, index, lookback, reset dates) are disclosed rather than assumed | Jan 2018 – May 2026 | 9.70 million loan-months; 253,376 loans |
| Credit — agency | Freddie Mac loan performance history is the estimation spine; the Fannie Mae panel enters only the pooled deep-delinquency refit and is otherwise scored as an untouched cross-agency comparison (§8) | 1999Q1 – 2026Q1 (deployed estimation panel 1999Q1 – 2025Q3) | — |
| Credit — government | Ginnie Mae monthly delinquency and removal records | Oct 2013 – Jul 2026 (154 consecutive disclosure months) | — |
Each span is a span of source reporting periods, not an evaluation window: §9 states where each book's evaluation panel ends. The agency credit row carries a third span in parentheses: the acquisition quarters the deployed conventional credit coefficients were estimated on, as recorded in the coefficient file itself. Those coefficients have not been re-estimated since the loan performance tables were extended, so the two spans in that row are different statements — the first about the data now held, the second about the fit — and /validation/credit states both. Freddie Mac completed-lifecycle histories additionally anchor the conventional curtailment and payoff-completion calibration. Fannie Mae and Freddie Mac collateral is treated as behaviorally equivalent by design.
Scope decisions in the estimation data
- Pre-2012 conventional observations are excluded. Refinance behavior in the HARP era was governed by program eligibility rules outside the model's characteristic set, and the pre-crisis underwriting regime lies outside the model's intended domain of application. Older vintages are retained solely for their seasoned post-2012 observations.
- A fixed observation cutoff (September 2025 for the conventional book) excludes the most recent reporting months, in which loans that have not yet reported constitute a substantial and non-random share of the population. The government fixed-rate estimation era framework closes at May 2026, so disclosure months after that date are ingested and evaluated but do not enter estimation.
- Estimation and evaluation populations are separated. Models are estimated on a fixed 45% loan sample; all acceptance decisions are measured on the full population. This separation is a standing control against over-fitting the estimation objectives.
Market and macroeconomic data
Beyond the loan-level performance records, the model consumes a deliberately small set of external market and macroeconomic series. Each is listed here with the role it plays; a series earns its place only by driving an identified behavioral mechanism, never as a free regressor.
- Primary mortgage rates. Freddie Mac Primary Mortgage Market Survey 30- and 15-year rates, monthly averages of weekly observations, January 2000 – August 2026, with no gaps. The government models use an FHA primary-rate series (January 2000 – August 2026); the 15-year government rate is derived from the conventional 30/15 spread, as no 15-year government survey exists (a documented approximation). Why it is included: the borrower-facing primary rate, set against the note rate, defines the refinance incentive — the model's central driver — and anchors the rate-history features (burnout, prior-opportunity effects) that condition it.
- Home prices. The FHFA all-transactions state-level house price index, quarterly, 1975–2026. Why it is included: it marks each loan's collateral to market, producing the current loan-to-value path that governs equity-driven prepayment (cash-out extraction, move-up turnover), refinance qualification, and the credit model's delinquency and default hazards. State-level resolution matters: equity positions of the same vintage differ materially across states.
- ARM indices. One-year constant-maturity Treasury (1953–2026) and the 30-day average SOFR (2018–2026), with the statutory LIBOR Act fallback spread applied after the June 2023 LIBOR cessation. Why they are included: adjustable-rate coupons reset off these indices; the model projects each loan's coupon path — and with it the payment shock and refinance incentive at each reset — from the contractual margin, caps, and the index path.
- Mortgage-industry employment. A monthly series of employment in the mortgage origination industry, driving a submodel of origination capacity (partial-adjustment specification, R² 0.83 over 2000–2026). Why it is included: when refinance demand surges, the industry's processing capacity — proxied by its staffing level, which adjusts with a lag — determines how fully in-the-money demand converts into closings. Identical incentives produce different realized speeds depending on how stretched originators are; this series is how the model knows.
- Refinance-eligible share of the market. A note-rate distribution simulation of the outstanding market, derived from the survey rate series rather than consumed externally (historical replay R² 0.949). Why it is included: media attention and originator marketing scale with the share of the market able to refinance, amplifying individual incentive — a loan's response depends partly on how much of the market is in the money alongside it.
- Housing turnover activity. Existing single-family home sales (seasonally adjusted annual rate) scaled by the size of the outstanding single-family mortgage universe. Why it is included: the housing-turnover channel — prepayment through property sale — moves with aggregate housing-market activity independently of rates; sales volume relative to the universe of mortgaged homes calibrates that baseline intensity through time.
- Conforming loan limits. The FHFA annual limit schedule. Why it is included: it defines the high-balance boundary of the conventional market, gating which loans face conforming refinance execution and which price into the jumbo market.
Two macroeconomic quantities deliberately enter as scenario inputs rather than fitted feeds. The forward home-price path is the user's assumption — the historical index grounds today's equity position; its future is a scenario. Unemployment likewise enters the credit model as a user-specified stress on delinquency formation — lagged twelve months and applied to rising unemployment only, reflecting the measured asymmetry of the response — rather than as a projected series. The model consumes no analyst forecast of any macroeconomic variable.
4The conventional model ↑ top
The conventional model is a structural multiplicative model of the post-2012 agency conforming universe. It is estimated by constrained numerical optimization against stratified cohort-level objectives — fit is measured simultaneously by period, geography, loan purpose, and along the incentive, age, burnout, credit score, loan-to-value, and balance dimensions — with observations weighted by unpaid principal balance. Its principal machinery:
- Refinance. The two-population threshold mixture with endogenous burnout described in §2, an amplification term active on sustained falling-rate paths (reflecting the observed media and solicitation response), and coupling to the capacity and eligibility submodels.
- Turnover. Level, seasonality, a seasoning profile, equity and rate lock-in responses, and the elevated-rate machinery that carries the 2022–25 turnover contraction. The contraction variables are structurally zero for all months before 2022, so they cannot contaminate the fit of earlier periods.
- Cash-out refinance. An incentive response gated by accumulated equity, with maximum activity at moderate equity levels rather than maximum equity, and a structural penalty when extraction requires refinancing the full balance to an above-market rate.
- Curtailment. A separately measured component with term-specific levels (seasoned curtailment is measured at 2.1% per annum on 15-year collateral against 0.7% on 30-year), a U-shaped seasoning profile, and diminished response at both deep in- and out-of-the-money extremes.
- Forward market submodels. Origination capacity (the 2020–21 refinance wave was throttled by industry capacity, and the model carries that constraint explicitly), the refinance-eligible share of the market, and a measure of turnover deferred during elevated-rate periods that is released as conditions normalize.
Representative-line projection uses the same engine: disclosed pool dispersion — coupon quartiles, state, purpose, occupancy, property type, servicer, and delinquency composition — is expanded into a weighted grid of constituent loans whose blended projection is required by automated test to agree with loan-by-loan reference computation (§10).
5Government fixed-rate models ↑ top
The FHA, VA, and USDA models are estimated on the Ginnie Mae loan-level disclosure over a common era framework (October 2013 – 2016, 2017–2019, 2020–2021, and 2022–2026) and share the conventional model's component architecture, with program-specific machinery where the programs genuinely differ:
- FHA. The mortgage insurance premium schedule as a dated cost ledger (annual premium 1.35% reduced to 0.85% effective January 2015 and to 0.55% effective March 2023, each tied to its Mortgagee Letter, with the 2015 reduction's refinance wave estimated as a policy event); the reduced-premium treatment of streamline refinances of pre-June-2009 endorsements; the 210-day streamline seasoning requirement; a delinquency-recency friction on refinance qualification; credit-score response; and a temporary-buydown expiry effect estimated from the matched +8 to +13 CPR differential observed as 2-1 buydown subsidies end.
- VA. The Interest Rate Reduction Refinance Loan (IRRRL) churn episode and its regulatory resolution: hazard rates on young IRRRLs fell from approximately 32 CPR in 2014–16 to under 3 CPR by 2019 as Ginnie Mae's pooling-eligibility memoranda (APM 16-05, APM 17-06) and the Economic Growth, Regulatory Relief, and Consumer Protection Act of 2018 took effect, each entering as a dated schedule; the seasoning, recoupment, and net-tangible-benefit requirements of 38 U.S.C. §3709; the 2019 cash-out loan-to-value restriction, including the surge of originations immediately preceding its effective date; the funding-fee schedule, including the 2020–23 increase; issuer-level speed differentials (67 issuers, estimated in four speed tiers, offered to users as reference values subject to their own judgment); and the foreclosure-moratorium and servicing-program policies of 2023–25.
- USDA. The smallest book (approximately 100 million loan-months), with program-faithful constraints — the program offers no cash-out refinance, so refinance activity is modeled accordingly — the annual-fee schedule, and a purchase-dominated seasonal profile.
All three models carry the temporary-buydown machinery (the VA and USDA fits are constrained to the shape of the better-identified FHA estimate), state-level factors only where a documented economic rationale exists, a common delinquency-transition convention, and coupling to the credit model (§8).
6Adjustable-rate models ↑ top
Agency ARM (v2.4)
Hybrid ARM prepayment is dominated by the reset schedule: anticipation ahead of the first reset, elevated activity in the reset month, an annual recurrence at subsequent resets, and a decaying elevation thereafter. The model estimates this structure on the Freddie Mac non-standard dataset. Two design decisions merit a validator's attention:
- The headline accuracy metric is defined on the investable population. Performance months before January 2004 constitute 21% of equally-weighted observation months but only 8% of balance-weighted exposure, and contribute 62% of total squared error, reflecting the small and volatile early book. The headline metric is therefore defined on performance months from January 2004 forward (92% of balance-weighted exposure); the full-history figure is reported alongside it as a diagnostic in all standard outputs. The estimation objective itself is unaffected by this reporting definition.
- Post-2021 behavior is a validated extrapolation, not an in-sample fit. The dataset's performance history ends in September 2021 and contains no rising-rate regime. The model's suppression of reset-month activity when all refinancing alternatives are prohibitively expensive is inactive in the estimation sample and was validated out of sample against 2022–25 Ginnie Mae FHA and VA ARM cohorts: the model's reset-month response, which initially overstated the matched observed response by a factor of four, reproduces it at a factor of 1.0 after the suppression term. A related moderation of activity on loans far from their first reset was estimated on the same government panel and applied to the agency model at the more conservative of the two program-level estimates, because assumable government loans can only understate the corresponding agency effect.
Contract caps are taken from disclosure where available; undisclosed caps default to the market-standard 5/2/5 convention, and the user may override them (see §11). Seller-level speed differentials and a re-inference procedure for implausible disclosed margins complete the specification.
Government ARM (v1.0)
A standalone model estimated on the Ginnie Mae ARM disclosure, in which the contract terms are disclosed rather than assumed — and are materially different from agency conventions: annual/annual/lifetime caps of 1/1/5 apply to 98.6% of the book, and 99.7% of loans reference one-year constant-maturity Treasury. The model's structure follows from these contracts. The coupon path is projected explicitly and feeds scheduled amortization. Under a 1%-per-year cap, a four-point rate shock resolves as a multi-year sequence of capped adjustments rather than a single event; consistent with this, the data show no reset-month spike, and the model contains no such term. The two agencies differ along two institutionally grounded dimensions, each estimated with the sign the institutional analysis predicts: assumption-driven retention is far deeper for FHA than VA (an FHA assumption requires a creditworthiness review, while a VA assumption requires substitution of entitlement), and the refinance escape from anticipated resets is approximately four times stronger for VA (the 0.50% minimum rate-reduction test of 38 U.S.C. §3709(b) applies only to fixed-to-fixed IRRRLs, leaving the ARM-to-fixed route open). The model is deliberately parsimonious — 41 scalar and 12 seasonal parameters, against roughly 110 for the government fixed-rate models — and was estimated through December 2024 with January 2025 – May 2026 withheld from estimation entirely.
7The Ginnie Mae II TBA composite ↑ top
A Ginnie Mae II TBA delivers a mixture of FHA, VA, and USDA collateral. The GNMA product projects each program through its own estimated model and combines the projections with survivorship-weighted mixture arithmetic under exact balance conservation, so that the program mix evolves over the projection as faster-paying collateral amortizes and prepays away. The composite introduces no additional parameters, and two identities are enforced by permanent automated test: a single-program pool reproduces that program's model output exactly, and the blended projection reproduces an independent loan-by-loan survivorship computation. The default mix is measured from current production — 2026Q1 originations: FHA 50.0%, VA 48.3%, USDA 1.7% by balance, from 366,752 loans totaling $126.6 billion — with a coupon-level reference table (VA reaches approximately 64% of the 5.0–5.5% coupon bucket). Pool dispersion runs through each program's own representative-line machinery, with program vocabulary translated appropriately, before blending. The composite models the collateral; it does not model the TBA delivery option.
8The credit model ↑ top
Delinquency and default are projected by a monthly state-transition model, coupled in both directions to the prepayment models. It is not a default curve applied after the fact. The object advanced from month to month is a distribution over payment states, and every credit quantity reported — delinquency stocks by severity, the default rate, the security-basis removal rate, and the delinquency conditioning applied to voluntary speeds — is a functional of that distribution.
The state machine and what its state carries
Each projected cell — a single loan, or one line of a representative-line grid — carries probability mass over the following states, advanced one calendar month at a time:
- Current, and the shallow ladder 30 / 60 / 90 days delinquent.
- Deep delinquency at exact months delinquent. Past ninety days the model does not hold a single bucket. It holds a separate state for every whole month delinquent from four through fifty-nine, the last absorbing. Cure rates, foreclosure timing, and repurchase eligibility all depend steeply on the delinquency counter itself, and only an exact-month representation can express a dated policy rule as what it is: a repurchase obligation attaching at twenty-four months delinquent is a state index here, not an estimated hazard.
- A parallel repeat ladder. A borrower entering delinquency from a previously delinquent state behaves measurably differently from a first-time entrant, and not in one direction: repeat episodes cure less readily at thirty days and more readily once deep. The two populations are therefore carried on separate ladder tracks rather than blended into one.
- Post-deep transit states. A loan that has just left deep delinquency for a shallow step behaves differently from an ordinary shallow-step loan for exactly one month; the measured grouping shows no memory beyond that, so these states have a one-month dwell by construction and then rejoin the repeat ladder.
- Two payment-history clocks. A cured borrower is not a never-delinquent borrower. Time since the episode is tracked month by month to three years, with a permanent state beyond, on two tracks distinguished by the worst status the episode reached — on the government books the two tracks are measured to be effectively indistinguishable and are carried at a common level (§11). Re-default risk is elevated immediately after cure and decays toward a permanent floor above the never-delinquent level.
The population is additionally represented as a two-mass mixture — a small subpopulation with a substantially elevated propensity to enter delinquency, alongside the remainder — because a homogeneous chain generates far more ever-delinquent borrowers over a loan's life than the book actually contains. With the mixture in place, the permanent re-default floor and the age profile of the share of borrowers carrying a prior delinquency emerge from the changing composition of the surviving pool rather than being fitted as levels.
The estimated transitions, and how a hazard is formed
Ten transitions are estimated and carried into the projection: entry (current to thirty days); a cure and a roll for each of the thirty-, sixty- and ninety-day steps; and three competing exits from deep delinquency — cure, voluntary payoff, and liquidation. Each is a monthly rate formed multiplicatively from four elements:
- a base monthly rate at the reference cell;
- continuous response curves on the loan's characteristics — credit score, loan-to-value, loan age, debt-to-income, months delinquent on the deep block, and, on the agency deep cure and liquidation hazards, original loan balance. The curves are smooth logistic-plus-ramp forms fitted through the estimated cell multipliers in log space, so each curve passes through the estimated cell mean at the right place and saturates outside the observed range rather than extrapolating a straight line into characteristics no loan in the estimation data displayed;
- a categorical seasonal factor; and
- a dated schedule of period effects whose windows are documented policy dates rather than fitted breakpoints — pre-crisis, the crisis, the recovery, the 2020 forbearance inflow, the 2021 forbearance exits, the 2022 foreclosure restart, and 2023 onward — with amplitudes estimated within those fixed windows.
Further hazards are estimated but deliberately excluded from the borrower projection: the enterprises' post-buyout disposal channel and the government issuers' delinquency-buyout hazard. Both are servicing and agency machinery rather than borrower behavior; they are estimated for description and for the dated policy overlay that produces the security basis, and are never projected as behavior (§2). An alternative entry specification driven by unemployment is likewise estimated and retained, but is not enabled.
The composition blocks are measured, not free
Around the ten hazards the engine carries a set of blocks that shape how mass moves. None of them is a knob. Each is a census of the loan-level panels, a documented policy schedule, or a level determined by a solve against measured targets:
| Block | What it carries |
|---|---|
| First-time versus repeat | The measured differential between first and repeat episodes at each shallow step, in both the cure and the roll direction |
| Cure destinations | A cure in the transition data is a move to any better state; the measured destination shares decide how much of it reaches current and how much is a partial improvement that continues the episode |
| Post-deep transit | The one-month multipliers applied to a loan that has just left deep delinquency |
| Payment history and the mixture | The two post-cure risk curves and the two-mass mixture. The relative risk of the elevated mass is pinned by the measured permanent re-default floor; its share and the shrinkage of the within-type transient are selected against the measured share of borrowers with a prior delinquency by loan age and against the measured delinquency occupancy; the never-delinquent entry level is solved to hold the measured aggregate delinquency level |
| Foreclosure timeline | The measured judicial versus non-judicial differential on deep liquidation (agency book only; §11) |
| Delinquency-counter dynamics | Roughly three-quarters of deep months advance the reported counter by one. A partial payment or a repayment plan holds it flat, and a small share reduces it. Measured on the agency panel census and common across the two enterprises |
| Roll destinations | The small measured share of thirty- and sixty-day rolls whose reported counter jumps two or more steps in one month, which a strict one-step ladder cannot produce |
| Deep-tail extension | The government books' deep ladder beyond the disclosure's counter cap (below) |
| Book composition by age | The measured composition of the current book by loan age, used to initialize pool projections (below) |
| Removal schedule | The security-basis translation (below) |
Where a book's coefficient artifact does not carry a block, the engine reduces exactly to the simpler model rather than substituting an assumed default. Taken together with the ten hazards, this means the engine forms no rate that is not either estimated from the panels, measured from a census, or determined by a solve against measured observations: there is no free parameter available to the projection. The only quantities a user supplies are the scenario inputs of §3 — the forward home-price path, and the unemployment stress on delinquency formation.
Probability mass is conserved. All state outflows are jointly bounded in the competing-risks sense, so no state can emit more probability than it holds, and steady-state delinquency compositions are fitted to measured targets recorded alongside each model's coefficients. On the agency book the deep-delinquency hazards carry dated period factors in common with the rest of the ladder; because their reference period is the estimation window itself, those factors act only where a loan's own history is replayed and leave every forward projection unchanged.
Modifications are treated as cures. A borrower current after modification is a cured borrower with elevated re-default risk, exactly as the loan-level data reports; servicer workout mechanics are not estimated as borrower behavior.
Reconstruction of borrower outcomes from behind servicing operations
On the government books, issuer buyouts remove 15–25% of deeply delinquent loans from pools each month, censoring observed outcomes; moreover, the commonly used liquidation file disagrees with the monthly delinquency records on the reason for roughly 46% of involuntary removals. The model therefore estimates liquidation hazards from the monthly records, extends the deep ladder beyond the disclosure's six-month cap (below), and assigns post-buyout outcomes using the published empirical literature on early-buyout loans — a probability of ultimate default given buyout of approximately 0.20, with a documented range of 0.17–0.30. One notable consequence of this correction: USDA carries the highest ultimate default rate of the three government programs; the lower uncorrected reading is an artifact of buyout censoring.
Estimation, and the government disclosure's six-month cap
The transitions are estimated as Poisson generalized linear models on full-population transition cubes built from the loan-level panels. At its optimum a Poisson model reproduces every included one-way margin of the data exactly, which serves as the convergence check.
On the agency book Freddie Mac is the estimation spine — it is the panel that spans the crisis — and the Fannie Mae panel is scored as an untouched cross-agency comparison rather than entering the fit. The deep-delinquency hazards are separately re-estimated on the pooled two-agency recent window under the single default definition described below, the two enterprises' collateral being treated as behaviorally equivalent by design (§3).
The government books are estimated on the Ginnie Mae loan-level disclosure, whose delinquency counter is capped at six months. Three consequences follow, and together they are the reason the government deep tail is constructed rather than simply fitted:
- Shape. Nothing past six months is observable as a counter, so the shape of the tail is carried from the agency book's exact-month profile, positioned so that the exposure-weighted average of the grafted tail over the six-months-and-beyond range reproduces the disclosure's aggregate for that bucket, while the two months the disclosure does identify individually — four and five — are reproduced exactly.
- Liquidation timing. For the liquidation hazard the tail is not borrowed. It is measured directly from the government disclosure's own record of how long a loan sits at the six-month cap and what fraction of that population leaves each month, deconvolved back onto the reported counter through the measured counter dynamics above. This identifies the foreclosure-timing profile past the cap from government data, and it is dated to the window on which it was measured.
- Level. The level of each beyond-cap exit is pinned rather than assumed. Cure and liquidation are solved jointly against the measured deep-delinquency stock, the measured aggregate delinquency level, and a follow-forward anchor for the eventual foreclosure rate of deeply delinquent loans. The beyond-cap voluntary payoff carries its own level, pinned by the program's own measured payoff rate out of deep delinquency, rather than inheriting the agency book's. Four unknowns are determined by four measured observations; the conditioning of that solve is recorded with the coefficients.
Seven products are covered by four estimated credit models. Conventional serves the conventional book and the agency ARM; FHA, VA and USDA serve their own books and the corresponding government ARM collateral; the Ginnie Mae II composite runs each constituent program through its own credit model and combines the results with the same survivorship mixture arithmetic used for prepayment (§7), introducing no additional credit parameters.
Coupling to the prepayment projection
The two models are not run side by side and added. The prepayment projection runs the clean path: delinquency is held at zero throughout the prepayment computation, and the credit engine owns the delinquency state. The credit engine consumes the projected voluntary speed as the rate at which its current pool leaves, and returns, for each month, the share of the surviving pool that is current at the start of that month — including cured borrowers — and the share that is delinquent. Each voluntary channel is then recomposed at those shares, and the total, the annualized speeds, and the channel decomposition are rebuilt from the recomposed parts.
The measured conditioning differs by book:
- Government books. Delinquent borrowers turn over at 2.3 to 2.9 times the current-borrower rate — distress and relocation sales — while refinance, cash-out and curtailment are effectively closed to them: a delinquent borrower cannot streamline, cannot qualify, and does not curtail.
- Agency book. Both turnover and refinance are suppressed among delinquent borrowers, to roughly 0.87 and 0.20 of the current-borrower rate respectively, while cash-out and curtailment carry no delinquency factor.
Because the channels move in different directions and their mix differs by program and by rate environment, the direction of the credit model's net effect on total voluntary speed is a consequence of that composition rather than an estimated correction. On adjustable-rate collateral the coupling runs one step further: the projected coupon path feeds delinquency formation, with a measured payment-shock elevation of entry for the year following an upward reset and a measured suppression for the two years following a downward one. The recomposition is applied once to an assembled projection, by construction.
Pools with undisclosed payment history are initialized at the composition the model itself generates for collateral of that age, obtained by replaying the loan's own history under the engine's dynamics, optionally conditioned on the disclosed share of borrowers with clean versus impaired pay histories; a security-basis book projection that carries its delinquent loans without a loan-level tape is primed instead from the measured current book's delinquency composition by loan age. On the government books that prior is the measured government book's own composition by program and loan age, so a government pool projection without a tape starts from the book being valued rather than from the model's stationary state.
Whole-loan and security valuation bases
The involuntary flow can be read on two bases, corresponding to two different owners of the cash flows.
- Loan basis — for whole-loan valuation. The involuntary flow is true default (CDR): the fitted liquidation hazard, with loss severity borne by the loan's owner. The default event is foreclosure completion — the acquisition of the property as real estate owned — recorded on a consistent basis across agencies; on the agency book the deep-delinquency hazards are estimated on the pooled Fannie Mae and Freddie Mac 2024–2025 panel under that single definition. Loan balance enters the deep-delinquency cure and liquidation hazards as a continuous covariate: at matched months delinquent and loan-to-value, small-balance loans liquidate at a higher rate and cure at a lower rate than large-balance loans, an agency-common borrower and servicing-economics effect measured on both panels. Servicer pool accounting is ignored, exactly as described above — a buyout is not a borrower event, and a whole-loan owner holds the loan through the workout.
- Security basis — for agency and government MBS. Any removal of a loan from a pool returns principal to the security holder at par under the guarantee, whether the removal is a delinquency buyout, a foreclosure, or a loss-mitigation purchase. From the bondholder's perspective these removals are prepayments: on this basis they constitute the involuntary speed — the Conditional Buyout Rate (CBR) — carry zero loss severity, and deplete the projected delinquency states — pools genuinely shed their delinquent loans, so pool-level delinquency stocks and the delinquency-conditioned CRR speeds remain internally consistent.
At the monthly-rate level the two bases share one identity: on the security (pool) basis, CPR = CRR + CBR; on the whole-loan (borrower) basis, CPR = CRR + CDR, with buyouts unwound to their true borrower outcome.
On the government books, security-basis removal rates are measured directly from the corrected monthly records by months delinquent, balance-weighted, over a trailing twelve-month window of the current policy regime. The measurement covers all three ways a delinquent loan leaves a pool at par — the delinquency buyout, the completed foreclosure, and the loss-mitigation purchase, the last of which carries the Department of Veterans Affairs' servicing purchase program on the VA book. Removal rates are negligible at thirty and sixty days and step up sharply at ninety, where pool buyout eligibility opens; past that the profile differs by program, FHA's peaking at ninety days as issuers exercise promptly while VA's and USDA's continue to rise with the delinquency counter. On the agency book no removal hazard is estimated at all: the enterprises' published delinquent-loan repurchase policy (since May 2020, repurchase at twenty-four months delinquent; earlier regimes differ) is applied as a dated rule to the model's exact-month delinquency ladder, together with the measured share of cures effected through modification — modification requires repurchase from the pool, while payment deferral does not — that share being measured on the loan's first modification, the repurchase event, rather than on the servicer's ever-modified flag. The removal schedule is a dated, current-policy quantity; historical regimes differ by an order of magnitude, and the schedule is re-measured at each data refresh rather than projected.
Validated at the observed delinquency composition of each government program's book, the security-basis removal flow (CBR) reproduces the observed involuntary removal rates of the most recent twelve months of disclosure to within half a percent of their level (FHA 1.617 projected versus 1.621 observed, VA 1.335 versus 1.342, USDA 0.945 versus 0.947, in annualized percent). Because the schedule is dated and the book's delinquency composition moves, these levels are re-measured and re-validated at each data refresh rather than held fixed; the corresponding figures for the preceding two-year window were materially lower. The loan basis is unchanged by the addition of the security basis, reproducing the previously validated model to the last representable digit. In integrations the basis defaults by instrument type — whole-loan analyses run the loan basis, pool and TBA analyses the security basis — and may be overridden per position.
9Performance summary ↑ top
The standing accuracy measure is balance-weighted monthly aggregate error in the voluntary prepayment series (CRR), stated in CPR percentage points, against realized history. It is computed on a standing evaluation universe rather than on the estimation sample: the full population for the government and adjustable-rate books, and the 56-vintage reference universe for the conventional book. Sub-period decompositions are tracked as first-class quantities. The figures below are the accepted values for the released versions and are regenerated from the same evaluation panels the validation report plots month by month; bias is reported as predicted minus actual.
| Book | CPR RMSE (pp) | Bias (pp) | Sub-period detail (bias unless noted) |
|---|---|---|---|
| Conventional | 1.51 | −0.01 | 2014–19 +0.09 · 2020–21 −0.00 · 2022–25 +0.04 · 2026 to date −0.09 |
| FHA | 1.06 | +0.06 | By period (RMSE/bias): 2013–16 0.95/+0.01 · 2017–19 1.04/+0.03 · 2020–21 1.16/+0.08 · 2022–25 0.80/−0.16 · 2026 to date 2.28/+2.03 |
| VA | 1.63 | +0.03 | By period (RMSE/bias): 2013–16 1.67/−0.04 · 2017–19 1.54/+0.03 · 2020–21 1.72/−0.00 · 2022–25 1.36/−0.25 · 2026 to date 2.83/+2.52 |
| USDA | 0.71 | −0.03 | By period (RMSE/bias): 2013–16 0.50/+0.11 · 2017–19 0.59/−0.32 · 2020–21 1.04/−0.20 · 2022–25 0.72/+0.11 · 2026 to date 0.78/+0.22 |
| ARM (investable, 2004– ) | 2.50 | −0.09 | 213 months, 92% of balance-weighted exposure; full-history diagnostic 3.48/−0.16 |
| Government ARM | 1.82 | +0.28 | 2018-01–2021-09 1.50/−0.04 · 2021-10–2022-08 2.96/−2.11 (weakest window) · 2022–25 2.17/+0.61 · 2026 to date 1.14/+0.07 · withheld window 2025-01–2026-05 2.03/+1.39 |
Government fixed-rate eras: Oct 2013–2016, 2017–2019, 2020–2021, 2022–2025, and 2026 to date. For every government book the 2022–2025 elevated-rate period and the 2026 months to date are reported separately and are never combined: a blended figure would net the elevated-rate under-prediction against the 2026 over-prediction. The conventional panel ends in March 2026, the government fixed-rate panels in July 2026, and the government ARM panel in May 2026. The government ARM window marked as withheld was excluded from estimation entirely. A standing development constraint requires the fixed-rate models not to under-predict the 2022–25 elevated-rate period; the VA figure of −0.25 and the FHA figure of −0.16 for that period are reviewed and accepted exceptions, each documented with its cause (§11).
For context on OPM's practicality as an interactive analytic: full-life single-loan projections complete in roughly 60–120 milliseconds per product, and a seven-scenario rate-shock analysis over a loan's full remaining life completes in under two-thirds of a second, on commodity hardware. These are wall-clock measurements; the benchmark suite that reproduces them accompanies the model source.
10Validation and controls ↑ top
Reconciliation of computational paths
The production system contains several optimized computational paths — batched representative-line grids, blended multi-state projections, and the Ginnie Mae II composite — that must produce the same results as straightforward loan-by-loan computation. Each such path is reconciled against its reference implementation by permanent automated test: raw monthly rates must agree to within the floating-point precision of the calculation, every displayed figure must agree at its displayed precision, single-program or single-state degenerate cases must reproduce the underlying model exactly, and blended projections must conserve balance exactly. A release cannot ship with any of these reconciliations failing.
Behavior-change detection
A fixed battery of representative requests — every product under multiple scenarios — is recorded and compared byte-for-byte across every change to the system. A change intended to be behavior-neutral must reproduce all recorded outputs identically; a change intended to alter model behavior must alter only the outputs it claims to alter, and the differences are quantified and reviewed before release (§12).
Scenario stress screening
A standing screening exercise projects several thousand actual loans — sampled to over-represent unusual occupancy, property type, term, credit, loan-to-value, and balance combinations — through a seven-scenario rate grid spanning −300 to +300 basis points over a ten-year horizon. Automated checks flag non-finite output, violations of scenario monotonicity, implausibly slow or fast projections, discontinuous month-to-month behavior, and inverted rally response. Flagged loans are reviewed rather than suppressed, and the identity of flagged loans — not merely their count — is compared between releases, so that offsetting changes cannot pass unnoticed.
Historical fit diagnostics and external benchmarks
Balance-weighted survival analysis compares modeled and realized loan survival across age, vintage, and characteristic bands, subject to minimum-exposure rules. Where independent external estimates exist, model responses are benchmarked against them; for example, the model's rate lock-in effect on housing turnover is compared with the Federal Housing Finance Agency's published estimate of approximately an 18% reduction in sale probability per percentage point of rate lock-in, estimated on approximately 50 million loans. The provenance of every external benchmark, including the period on which it was estimated, is recorded (§11).
Regression testing and independent review
The automated regression suite currently comprises approximately 1,000 tests. Each test written to guard a specific defect is demonstrated to fail against the defective code before the fix is accepted, so the suite's protective value is verified rather than presumed. In August 2026 the codebase underwent a structured internal review: ten independent area-by-area examinations produced 48 candidate findings; each was then independently re-examined with runtime reproduction, confirming 45 and refuting 3; all 45 confirmed findings were resolved, with the 17 that affected accepted model output corrected through the quantified-impact review process of §12, one reversible change at a time. The complete findings ledger, including reproduction evidence for each item, is retained as part of the model documentation.
Out-of-sample protocol
A documented out-of-sample protocol specifies two temporal splits (estimation through 2018 with evaluation on 2019–2021, and estimation through 2021 with evaluation on 2022–2025), acceptance criteria (absolute evaluation-period bias within approximately 0.5 CPR points; no incentive cohort deteriorating by more than 2 points), and — deliberately — a register of the protocol's own information-leakage channels, such as external benchmarks published from evaluation-period data. All validation is conditional: realized rates and home prices are supplied, and the behavioral response is what is scored. Execution of the full protocol is a planned exercise; its design and leakage register form part of this documentation set.
11Limitations ↑ top
The suite's documentation practice is to state residual weaknesses explicitly rather than absorb them into the fit. The material limitations follow.
- Conditional validation. All accuracy figures condition on realized rate and home-price paths. No claim is made regarding rate forecasting.
- Domain of application. The conventional model is scoped to the post-2012 agency conforming universe. Pre-crisis underwriting regimes and HARP-style eligibility programs lie outside its domain.
- ARM extrapolation. The agency ARM estimation data end in September 2021 and contain no rising-rate regime; post-2021 reset behavior is an extrapolation validated on government ARM collateral (§6) — a cross-collateral transfer, conservatively applied, but a transfer nonetheless. Undisclosed contract caps default to the 5/2/5 convention; the government contract audit demonstrates that convention does not describe government collateral, and disclosed caps with a user override are the mitigation. The government ARM panel contains no rally and no pre-2018 history; its weakest window is October 2021 – August 2022 (RMSE 2.96, bias −2.11); and margins are imputed for 30.5% of balance from an estimated program-by-vintage table.
- Forward index paths. Projected ARM indices hold their final observed fixing plus the scenario shock; a full short-rate term structure is not currently a model input. This principally affects ARM reset projections under steep-curve scenarios.
- Forward submodel uncertainty. The origination-capacity submodel's out-of-sample error (estimation restricted to pre-2018 data) is approximately 2.5 times its in-sample error at a three-year horizon; the forward capacity path should be treated as a scenario input with a wide band, not a point forecast.
- Credit censoring and literature transfer. Government deep-delinquency outcomes are buyout-censored, and the reconstruction relies in part on a published early-buyout study whose sample differs from the loans to which it is applied; the documented transfer range is retained. Agency liquidation hazards are right-censored by GSE note sales. Post-delinquency risk states are untyped on the government books by measured decision: a full-population re-measurement of the Ginnie panels found that episode depth adds little re-default signal beyond episode recency there (roughly a 1.05–1.14× differential in the first months after cure, against 1.46× on the agency book), with the deep-episode sample additionally thinned by buyout censoring; the distinction is therefore excluded rather than weakly fitted. Judicial-versus- non-judicial liquidation timing is modeled on the agency book only.
- Documented residuals by book. Each model maintains an explicit register of accepted residual errors. Notable entries: the FHA delinquency-severity curve is constant across eras, deferring a forbearance-regime variable (residuals within the affected cells reach +10 CPR in the 2020+ period); the VA model over-predicts moderately out-of-the-money refinance activity by +1.7 to +3.2 CPR, a residual retained after extensive candidate testing because every correction attempted degraded era-level bias; and the USDA model accepts a U.S.-territories residual and several artifacts in cells with negligible exposure. The VA figure for the 2022–25 period (−0.25) reflects a specific reviewed exception: a correction to the temporary-buydown machinery exposed a pre-existing under-prediction that had previously been offset, and the exposure was accepted and documented rather than re-concealed; the 2026 months to date (+2.52) are reported separately and are not netted against it. The FHA figure for the 2022–25 period (−0.16) reflects a reviewed input correction: the refinance-eligible share series the FHA model consumes is now computed under the FHA model's own rate-lag structure rather than the conventional model's; the coefficients were not re-estimated, the published figures were re-based to the corrected input, and the resulting deepening of the 2022–25 under-prediction (from −0.14) was accepted and documented rather than offset.
- Structural parameters. Thirty-two of the conventional model's 250 registered parameters are set structurally rather than estimated, each flagged in the parameter registry; the estimation procedure cannot alter them.
- Independence assumption in representative-line mode. Disclosed pool marginal distributions are treated as independent, as joint distributions are not disclosed at pool level. The reconciliation controls of §10 bound the mechanics; the assumption itself is irreducible from pool-level disclosure.
- Government 15-year rates. Derived from the conventional 30/15 spread in the absence of a government 15-year survey.
12Governance and change control ↑ top
- Versioned, self-describing model artifacts. Every model ships as a versioned coefficient artifact recording its estimation data inventory, sample counts, fit provenance, steady-state targets, and caveats. The deployed application reports the live version of every model on request, so any forecast is attributable to an exact build.
- Quantified-impact review. No change that alters accepted model output is released silently. Each candidate change is presented with its measured effect on every standing accuracy measure — headline errors, sub-period decompositions, and the stress-screening results — and is accepted or declined deliberately by the model owner. Accepted exceptions, such as the VA item in §11, are recorded as such.
- Controls on every change. The behavior-change detection battery, the computational reconciliations, the stress-screening comparisons, and the full regression suite run across every change: work intended to be neutral must prove neutrality, and work intended to change behavior must quantify itself.
- Periodic independent review. The August 2026 review (§10) — independent examination, adversarial verification with runtime reproduction, and fully quantified resolution — is the template for recurring review of the codebase.
This paper describes the Octant Prepay Model (OPM), the modeling suite behind the forecaster at octantanalytics.com. It is model documentation, not investment advice; reported figures describe historical fit and validation evidence, not guarantees of future performance. Detailed behavioral documentation, policy catalogs, the review ledger, and the out-of-sample protocol are maintained alongside the model source and are available to validation teams on request.