Octant Analytics Whitepaper

Model documentation

The Octant Prepay Model:
Methodology & Validation

The Octant Prepay Model (OPM) is a structural, loan-level prepayment and credit modeling framework covering seven agency and government mortgage products. This document describes OPM's modeling approach, data sources, estimation, validation evidence, known limitations, and change-control practices, at the level of detail required to scope a formal model validation.

Document date September 2026  ·  Model versions Conventional v89.86_bal30frz · FHA v1.25 · VA v2.19 · USDA v1.10 · ARM v2.4 · Government ARM v1.0 · GNMA composite
Live demonstration octantanalytics.com — the version identifiers above name the builds behind the accuracy figures in §9 and behind the validation report. The application's build fingerprint is an operator-only field and is not returned with a forecast, so a result obtained from the public application is not self-identifying.  ·  Companion the validation report — predicted vs actual speeds, month by month, for every book.

1Purpose and scope

OPM produces monthly voluntary prepayment projections (SMM and CPR — Conditional Prepayment Rate, the total prepayment speed; the voluntary component alone is the Conditional Repayment Rate, CRR) and credit projections (delinquency transitions and involuntary liquidation — the Conditional Default Rate, CDR) for U.S. residential mortgage collateral, at the single-loan, representative-line, and portfolio level, under user-specified interest-rate, home-price, and market scenarios. Seven products are covered, spanning both agency collateral (Fannie Mae and Freddie Mac) and government collateral (Ginnie Mae):

ProductCollateralVersionEstimation data
ConventionalAgency conforming fixed-rate, post-2012 (UMBS TBA)v89.86_bal30frzFannie Mae single-family loan performance
FHAGovernment fixed-rate, FHA-insuredv1.25Ginnie Mae monthly loan-level disclosure
VAGovernment fixed-rate, VA-guaranteedv2.19
USDAGovernment fixed-rate, USDA/RHS-guaranteedv1.10
ARMAgency hybrid adjustable-rate (3/1 through 10/1)v2.4Freddie Mac non-standard dataset
Government ARMGinnie Mae FHA and VA hybrid adjustable-ratev1.0Ginnie Mae ARM disclosure panel
GNMA (G2 TBA)Ginnie Mae II mixed FHA/VA/USDA collateralcompositeNo additional parameters — see §7

The models are conditional: given a realized or assumed path of rates and home prices, they project prepayment and credit behavior. They do not forecast interest rates. Validation accordingly measures the accuracy of the behavioral response against realized history, not the foresight of any scenario.

Four deployment surfaces share a single computational core: an interactive single-loan forecaster (this site); a representative-line mode that expands disclosed pool characteristics into a weighted grid of constituent loans; a portfolio scenario engine; and an adapter implementing a third-party analytics platform's prepayment and loss function interface.

2Modeling approach ↑ top

Structural decomposition

Each fixed-rate model decomposes total voluntary prepayment (CRR) into four economically distinct components: housing turnover (present in all environments and only weakly rate-sensitive), rate refinance (driven by the borrower's rate incentive and gated by qualification), cash-out refinance (gated by accumulated equity, and structurally penalized when extracting equity requires resetting the full balance to an above-market rate), and curtailment (partial principal prepayments, which are income-driven and exhibit diminished response at both deep rate extremes). The reported decomposition reconciles exactly to the reported total in every month. This structure is the principal defense against regime confounding: an estimated effect must reside on the component whose economics produce it.

Constrained, interpretable parameterization

Components are constructed from parametric response curves on continuous characteristics (incentive, loan age, credit score, loan-to-value, balance), categorical multipliers, and seasoning profiles. Sign and range restrictions are embedded in the parameterization itself rather than imposed as penalties, so that economically inadmissible states — a negative maximum turnover rate, an inverted refinance response — are unreachable by the estimation procedure. The refinance component is specified as a mixture of two borrower populations with differing refinance thresholds; a loan that remains outstanding despite a persistent rate incentive is progressively re-classified toward the less responsive population. This makes refinance burnout an endogenous property of the specification and the incentive response structurally monotone. Of the conventional model's 250 registered parameters, 218 are estimated and 32 are structural values set by the modeler where the estimation objective cannot identify them; each of the latter is individually flagged in the parameter registry.

Forward-looking regime variables

Regime effects — the 2020–21 refinance wave, the 2022–25 turnover contraction — are driven exclusively by functions of the interest-rate path, such as the current rate relative to its trailing average and the duration of an elevated-rate state. Calendar and vintage indicator variables are excluded by design: an indicator fits history but conveys no information under a forward scenario, whereas a rate-path variable generalizes to any scenario by construction.

Program rules as dated policy, not estimated behavior

Government program rules — FHA mortgage insurance premium schedules, the VA seasoning, recoupment, and net-tangible-benefit requirements of 38 U.S.C. §3709, streamline refinance eligibility windows, funding-fee changes — enter the models as dated policy schedules with effective dates taken from the primary sources (Mortgagee Letters, All-Participants Memoranda, public laws), rather than as estimated coefficients on calendar time. Conversely, issuer and servicer operations (such as delinquent-loan buyout timing) are never estimated as borrower behavior; the credit model explicitly reconstructs borrower outcomes from behind buyout censoring (§8).

3Data ↑ top

Loan performance data

BookSourceObservation spanScale
ConventionalFannie Mae single-family loan performance files, 56 acquisition-vintage files, 2011Q4–2025Q3Jan 2012 – Sep 20251.67 billion loan-months; 16.5 million payoff events
FHA / VA / USDAGinnie Mae monthly loan-level disclosure (Ginnie Mae I and II)Oct 2013 – Jul 2026 (154 consecutive disclosure months)FHA segment 1.08 billion loan-months
ARMFreddie Mac non-standard dataset, hybrid ARMs 3/1–10/1 (6.4 million origination records)Mar 1999 – Sep 2021 (performance history ends at the vendor's cutoff)81.7 million loan-months
Government ARMGinnie Mae ARM disclosure; contract terms (caps, index, lookback, reset dates) are disclosed rather than assumedJan 2018 – May 20269.70 million loan-months; 253,376 loans
Credit — agencyFreddie Mac loan performance history is the estimation spine; the Fannie Mae panel enters only the pooled deep-delinquency refit and is otherwise scored as an untouched cross-agency comparison (§8)1999Q1 – 2026Q1 (deployed estimation panel 1999Q1 – 2025Q3)
Credit — governmentGinnie Mae monthly delinquency and removal recordsOct 2013 – Jul 2026 (154 consecutive disclosure months)

Each span is a span of source reporting periods, not an evaluation window: §9 states where each book's evaluation panel ends. The agency credit row carries a third span in parentheses: the acquisition quarters the deployed conventional credit coefficients were estimated on, as recorded in the coefficient file itself. Those coefficients have not been re-estimated since the loan performance tables were extended, so the two spans in that row are different statements — the first about the data now held, the second about the fit — and /validation/credit states both. Freddie Mac completed-lifecycle histories additionally anchor the conventional curtailment and payoff-completion calibration. Fannie Mae and Freddie Mac collateral is treated as behaviorally equivalent by design.

Scope decisions in the estimation data

Market and macroeconomic data

Beyond the loan-level performance records, the model consumes a deliberately small set of external market and macroeconomic series. Each is listed here with the role it plays; a series earns its place only by driving an identified behavioral mechanism, never as a free regressor.

Two macroeconomic quantities deliberately enter as scenario inputs rather than fitted feeds. The forward home-price path is the user's assumption — the historical index grounds today's equity position; its future is a scenario. Unemployment likewise enters the credit model as a user-specified stress on delinquency formation — lagged twelve months and applied to rising unemployment only, reflecting the measured asymmetry of the response — rather than as a projected series. The model consumes no analyst forecast of any macroeconomic variable.

4The conventional model ↑ top

The conventional model is a structural multiplicative model of the post-2012 agency conforming universe. It is estimated by constrained numerical optimization against stratified cohort-level objectives — fit is measured simultaneously by period, geography, loan purpose, and along the incentive, age, burnout, credit score, loan-to-value, and balance dimensions — with observations weighted by unpaid principal balance. Its principal machinery:

Representative-line projection uses the same engine: disclosed pool dispersion — coupon quartiles, state, purpose, occupancy, property type, servicer, and delinquency composition — is expanded into a weighted grid of constituent loans whose blended projection is required by automated test to agree with loan-by-loan reference computation (§10).

5Government fixed-rate models ↑ top

The FHA, VA, and USDA models are estimated on the Ginnie Mae loan-level disclosure over a common era framework (October 2013 – 2016, 2017–2019, 2020–2021, and 2022–2026) and share the conventional model's component architecture, with program-specific machinery where the programs genuinely differ:

All three models carry the temporary-buydown machinery (the VA and USDA fits are constrained to the shape of the better-identified FHA estimate), state-level factors only where a documented economic rationale exists, a common delinquency-transition convention, and coupling to the credit model (§8).

6Adjustable-rate models ↑ top

Agency ARM (v2.4)

Hybrid ARM prepayment is dominated by the reset schedule: anticipation ahead of the first reset, elevated activity in the reset month, an annual recurrence at subsequent resets, and a decaying elevation thereafter. The model estimates this structure on the Freddie Mac non-standard dataset. Two design decisions merit a validator's attention:

Contract caps are taken from disclosure where available; undisclosed caps default to the market-standard 5/2/5 convention, and the user may override them (see §11). Seller-level speed differentials and a re-inference procedure for implausible disclosed margins complete the specification.

Government ARM (v1.0)

A standalone model estimated on the Ginnie Mae ARM disclosure, in which the contract terms are disclosed rather than assumed — and are materially different from agency conventions: annual/annual/lifetime caps of 1/1/5 apply to 98.6% of the book, and 99.7% of loans reference one-year constant-maturity Treasury. The model's structure follows from these contracts. The coupon path is projected explicitly and feeds scheduled amortization. Under a 1%-per-year cap, a four-point rate shock resolves as a multi-year sequence of capped adjustments rather than a single event; consistent with this, the data show no reset-month spike, and the model contains no such term. The two agencies differ along two institutionally grounded dimensions, each estimated with the sign the institutional analysis predicts: assumption-driven retention is far deeper for FHA than VA (an FHA assumption requires a creditworthiness review, while a VA assumption requires substitution of entitlement), and the refinance escape from anticipated resets is approximately four times stronger for VA (the 0.50% minimum rate-reduction test of 38 U.S.C. §3709(b) applies only to fixed-to-fixed IRRRLs, leaving the ARM-to-fixed route open). The model is deliberately parsimonious — 41 scalar and 12 seasonal parameters, against roughly 110 for the government fixed-rate models — and was estimated through December 2024 with January 2025 – May 2026 withheld from estimation entirely.

7The Ginnie Mae II TBA composite ↑ top

A Ginnie Mae II TBA delivers a mixture of FHA, VA, and USDA collateral. The GNMA product projects each program through its own estimated model and combines the projections with survivorship-weighted mixture arithmetic under exact balance conservation, so that the program mix evolves over the projection as faster-paying collateral amortizes and prepays away. The composite introduces no additional parameters, and two identities are enforced by permanent automated test: a single-program pool reproduces that program's model output exactly, and the blended projection reproduces an independent loan-by-loan survivorship computation. The default mix is measured from current production — 2026Q1 originations: FHA 50.0%, VA 48.3%, USDA 1.7% by balance, from 366,752 loans totaling $126.6 billion — with a coupon-level reference table (VA reaches approximately 64% of the 5.0–5.5% coupon bucket). Pool dispersion runs through each program's own representative-line machinery, with program vocabulary translated appropriately, before blending. The composite models the collateral; it does not model the TBA delivery option.

8The credit model ↑ top

Delinquency and default are projected by a monthly state-transition model, coupled in both directions to the prepayment models. It is not a default curve applied after the fact. The object advanced from month to month is a distribution over payment states, and every credit quantity reported — delinquency stocks by severity, the default rate, the security-basis removal rate, and the delinquency conditioning applied to voluntary speeds — is a functional of that distribution.

The state machine and what its state carries

Each projected cell — a single loan, or one line of a representative-line grid — carries probability mass over the following states, advanced one calendar month at a time:

The population is additionally represented as a two-mass mixture — a small subpopulation with a substantially elevated propensity to enter delinquency, alongside the remainder — because a homogeneous chain generates far more ever-delinquent borrowers over a loan's life than the book actually contains. With the mixture in place, the permanent re-default floor and the age profile of the share of borrowers carrying a prior delinquency emerge from the changing composition of the surviving pool rather than being fitted as levels.

The estimated transitions, and how a hazard is formed

Ten transitions are estimated and carried into the projection: entry (current to thirty days); a cure and a roll for each of the thirty-, sixty- and ninety-day steps; and three competing exits from deep delinquency — cure, voluntary payoff, and liquidation. Each is a monthly rate formed multiplicatively from four elements:

Further hazards are estimated but deliberately excluded from the borrower projection: the enterprises' post-buyout disposal channel and the government issuers' delinquency-buyout hazard. Both are servicing and agency machinery rather than borrower behavior; they are estimated for description and for the dated policy overlay that produces the security basis, and are never projected as behavior (§2). An alternative entry specification driven by unemployment is likewise estimated and retained, but is not enabled.

The composition blocks are measured, not free

Around the ten hazards the engine carries a set of blocks that shape how mass moves. None of them is a knob. Each is a census of the loan-level panels, a documented policy schedule, or a level determined by a solve against measured targets:

BlockWhat it carries
First-time versus repeatThe measured differential between first and repeat episodes at each shallow step, in both the cure and the roll direction
Cure destinationsA cure in the transition data is a move to any better state; the measured destination shares decide how much of it reaches current and how much is a partial improvement that continues the episode
Post-deep transitThe one-month multipliers applied to a loan that has just left deep delinquency
Payment history and the mixtureThe two post-cure risk curves and the two-mass mixture. The relative risk of the elevated mass is pinned by the measured permanent re-default floor; its share and the shrinkage of the within-type transient are selected against the measured share of borrowers with a prior delinquency by loan age and against the measured delinquency occupancy; the never-delinquent entry level is solved to hold the measured aggregate delinquency level
Foreclosure timelineThe measured judicial versus non-judicial differential on deep liquidation (agency book only; §11)
Delinquency-counter dynamicsRoughly three-quarters of deep months advance the reported counter by one. A partial payment or a repayment plan holds it flat, and a small share reduces it. Measured on the agency panel census and common across the two enterprises
Roll destinationsThe small measured share of thirty- and sixty-day rolls whose reported counter jumps two or more steps in one month, which a strict one-step ladder cannot produce
Deep-tail extensionThe government books' deep ladder beyond the disclosure's counter cap (below)
Book composition by ageThe measured composition of the current book by loan age, used to initialize pool projections (below)
Removal scheduleThe security-basis translation (below)

Where a book's coefficient artifact does not carry a block, the engine reduces exactly to the simpler model rather than substituting an assumed default. Taken together with the ten hazards, this means the engine forms no rate that is not either estimated from the panels, measured from a census, or determined by a solve against measured observations: there is no free parameter available to the projection. The only quantities a user supplies are the scenario inputs of §3 — the forward home-price path, and the unemployment stress on delinquency formation.

Probability mass is conserved. All state outflows are jointly bounded in the competing-risks sense, so no state can emit more probability than it holds, and steady-state delinquency compositions are fitted to measured targets recorded alongside each model's coefficients. On the agency book the deep-delinquency hazards carry dated period factors in common with the rest of the ladder; because their reference period is the estimation window itself, those factors act only where a loan's own history is replayed and leave every forward projection unchanged.

Modifications are treated as cures. A borrower current after modification is a cured borrower with elevated re-default risk, exactly as the loan-level data reports; servicer workout mechanics are not estimated as borrower behavior.

Reconstruction of borrower outcomes from behind servicing operations

On the government books, issuer buyouts remove 15–25% of deeply delinquent loans from pools each month, censoring observed outcomes; moreover, the commonly used liquidation file disagrees with the monthly delinquency records on the reason for roughly 46% of involuntary removals. The model therefore estimates liquidation hazards from the monthly records, extends the deep ladder beyond the disclosure's six-month cap (below), and assigns post-buyout outcomes using the published empirical literature on early-buyout loans — a probability of ultimate default given buyout of approximately 0.20, with a documented range of 0.17–0.30. One notable consequence of this correction: USDA carries the highest ultimate default rate of the three government programs; the lower uncorrected reading is an artifact of buyout censoring.

Estimation, and the government disclosure's six-month cap

The transitions are estimated as Poisson generalized linear models on full-population transition cubes built from the loan-level panels. At its optimum a Poisson model reproduces every included one-way margin of the data exactly, which serves as the convergence check.

On the agency book Freddie Mac is the estimation spine — it is the panel that spans the crisis — and the Fannie Mae panel is scored as an untouched cross-agency comparison rather than entering the fit. The deep-delinquency hazards are separately re-estimated on the pooled two-agency recent window under the single default definition described below, the two enterprises' collateral being treated as behaviorally equivalent by design (§3).

The government books are estimated on the Ginnie Mae loan-level disclosure, whose delinquency counter is capped at six months. Three consequences follow, and together they are the reason the government deep tail is constructed rather than simply fitted:

Seven products are covered by four estimated credit models. Conventional serves the conventional book and the agency ARM; FHA, VA and USDA serve their own books and the corresponding government ARM collateral; the Ginnie Mae II composite runs each constituent program through its own credit model and combines the results with the same survivorship mixture arithmetic used for prepayment (§7), introducing no additional credit parameters.

Coupling to the prepayment projection

The two models are not run side by side and added. The prepayment projection runs the clean path: delinquency is held at zero throughout the prepayment computation, and the credit engine owns the delinquency state. The credit engine consumes the projected voluntary speed as the rate at which its current pool leaves, and returns, for each month, the share of the surviving pool that is current at the start of that month — including cured borrowers — and the share that is delinquent. Each voluntary channel is then recomposed at those shares, and the total, the annualized speeds, and the channel decomposition are rebuilt from the recomposed parts.

The measured conditioning differs by book:

Because the channels move in different directions and their mix differs by program and by rate environment, the direction of the credit model's net effect on total voluntary speed is a consequence of that composition rather than an estimated correction. On adjustable-rate collateral the coupling runs one step further: the projected coupon path feeds delinquency formation, with a measured payment-shock elevation of entry for the year following an upward reset and a measured suppression for the two years following a downward one. The recomposition is applied once to an assembled projection, by construction.

Pools with undisclosed payment history are initialized at the composition the model itself generates for collateral of that age, obtained by replaying the loan's own history under the engine's dynamics, optionally conditioned on the disclosed share of borrowers with clean versus impaired pay histories; a security-basis book projection that carries its delinquent loans without a loan-level tape is primed instead from the measured current book's delinquency composition by loan age. On the government books that prior is the measured government book's own composition by program and loan age, so a government pool projection without a tape starts from the book being valued rather than from the model's stationary state.

Whole-loan and security valuation bases

The involuntary flow can be read on two bases, corresponding to two different owners of the cash flows.

At the monthly-rate level the two bases share one identity: on the security (pool) basis, CPR = CRR + CBR; on the whole-loan (borrower) basis, CPR = CRR + CDR, with buyouts unwound to their true borrower outcome.

On the government books, security-basis removal rates are measured directly from the corrected monthly records by months delinquent, balance-weighted, over a trailing twelve-month window of the current policy regime. The measurement covers all three ways a delinquent loan leaves a pool at par — the delinquency buyout, the completed foreclosure, and the loss-mitigation purchase, the last of which carries the Department of Veterans Affairs' servicing purchase program on the VA book. Removal rates are negligible at thirty and sixty days and step up sharply at ninety, where pool buyout eligibility opens; past that the profile differs by program, FHA's peaking at ninety days as issuers exercise promptly while VA's and USDA's continue to rise with the delinquency counter. On the agency book no removal hazard is estimated at all: the enterprises' published delinquent-loan repurchase policy (since May 2020, repurchase at twenty-four months delinquent; earlier regimes differ) is applied as a dated rule to the model's exact-month delinquency ladder, together with the measured share of cures effected through modification — modification requires repurchase from the pool, while payment deferral does not — that share being measured on the loan's first modification, the repurchase event, rather than on the servicer's ever-modified flag. The removal schedule is a dated, current-policy quantity; historical regimes differ by an order of magnitude, and the schedule is re-measured at each data refresh rather than projected.

Validated at the observed delinquency composition of each government program's book, the security-basis removal flow (CBR) reproduces the observed involuntary removal rates of the most recent twelve months of disclosure to within half a percent of their level (FHA 1.617 projected versus 1.621 observed, VA 1.335 versus 1.342, USDA 0.945 versus 0.947, in annualized percent). Because the schedule is dated and the book's delinquency composition moves, these levels are re-measured and re-validated at each data refresh rather than held fixed; the corresponding figures for the preceding two-year window were materially lower. The loan basis is unchanged by the addition of the security basis, reproducing the previously validated model to the last representable digit. In integrations the basis defaults by instrument type — whole-loan analyses run the loan basis, pool and TBA analyses the security basis — and may be overridden per position.

9Performance summary ↑ top

The standing accuracy measure is balance-weighted monthly aggregate error in the voluntary prepayment series (CRR), stated in CPR percentage points, against realized history. It is computed on a standing evaluation universe rather than on the estimation sample: the full population for the government and adjustable-rate books, and the 56-vintage reference universe for the conventional book. Sub-period decompositions are tracked as first-class quantities. The figures below are the accepted values for the released versions and are regenerated from the same evaluation panels the validation report plots month by month; bias is reported as predicted minus actual.

BookCPR RMSE (pp)Bias (pp)Sub-period detail (bias unless noted)
Conventional1.51−0.012014–19 +0.09 · 2020–21 −0.00 · 2022–25 +0.04 · 2026 to date −0.09
FHA1.06+0.06By period (RMSE/bias): 2013–16 0.95/+0.01 · 2017–19 1.04/+0.03 · 2020–21 1.16/+0.08 · 2022–25 0.80/−0.16 · 2026 to date 2.28/+2.03
VA1.63+0.03By period (RMSE/bias): 2013–16 1.67/−0.04 · 2017–19 1.54/+0.03 · 2020–21 1.72/−0.00 · 2022–25 1.36/−0.25 · 2026 to date 2.83/+2.52
USDA0.71−0.03By period (RMSE/bias): 2013–16 0.50/+0.11 · 2017–19 0.59/−0.32 · 2020–21 1.04/−0.20 · 2022–25 0.72/+0.11 · 2026 to date 0.78/+0.22
ARM (investable, 2004– )2.50−0.09213 months, 92% of balance-weighted exposure; full-history diagnostic 3.48/−0.16
Government ARM1.82+0.282018-01–2021-09 1.50/−0.04 · 2021-10–2022-08 2.96/−2.11 (weakest window) · 2022–25 2.17/+0.61 · 2026 to date 1.14/+0.07 · withheld window 2025-01–2026-05 2.03/+1.39

Government fixed-rate eras: Oct 2013–2016, 2017–2019, 2020–2021, 2022–2025, and 2026 to date. For every government book the 2022–2025 elevated-rate period and the 2026 months to date are reported separately and are never combined: a blended figure would net the elevated-rate under-prediction against the 2026 over-prediction. The conventional panel ends in March 2026, the government fixed-rate panels in July 2026, and the government ARM panel in May 2026. The government ARM window marked as withheld was excluded from estimation entirely. A standing development constraint requires the fixed-rate models not to under-predict the 2022–25 elevated-rate period; the VA figure of −0.25 and the FHA figure of −0.16 for that period are reviewed and accepted exceptions, each documented with its cause (§11).

For context on OPM's practicality as an interactive analytic: full-life single-loan projections complete in roughly 60–120 milliseconds per product, and a seven-scenario rate-shock analysis over a loan's full remaining life completes in under two-thirds of a second, on commodity hardware. These are wall-clock measurements; the benchmark suite that reproduces them accompanies the model source.

10Validation and controls ↑ top

Reconciliation of computational paths

The production system contains several optimized computational paths — batched representative-line grids, blended multi-state projections, and the Ginnie Mae II composite — that must produce the same results as straightforward loan-by-loan computation. Each such path is reconciled against its reference implementation by permanent automated test: raw monthly rates must agree to within the floating-point precision of the calculation, every displayed figure must agree at its displayed precision, single-program or single-state degenerate cases must reproduce the underlying model exactly, and blended projections must conserve balance exactly. A release cannot ship with any of these reconciliations failing.

Behavior-change detection

A fixed battery of representative requests — every product under multiple scenarios — is recorded and compared byte-for-byte across every change to the system. A change intended to be behavior-neutral must reproduce all recorded outputs identically; a change intended to alter model behavior must alter only the outputs it claims to alter, and the differences are quantified and reviewed before release (§12).

Scenario stress screening

A standing screening exercise projects several thousand actual loans — sampled to over-represent unusual occupancy, property type, term, credit, loan-to-value, and balance combinations — through a seven-scenario rate grid spanning −300 to +300 basis points over a ten-year horizon. Automated checks flag non-finite output, violations of scenario monotonicity, implausibly slow or fast projections, discontinuous month-to-month behavior, and inverted rally response. Flagged loans are reviewed rather than suppressed, and the identity of flagged loans — not merely their count — is compared between releases, so that offsetting changes cannot pass unnoticed.

Historical fit diagnostics and external benchmarks

Balance-weighted survival analysis compares modeled and realized loan survival across age, vintage, and characteristic bands, subject to minimum-exposure rules. Where independent external estimates exist, model responses are benchmarked against them; for example, the model's rate lock-in effect on housing turnover is compared with the Federal Housing Finance Agency's published estimate of approximately an 18% reduction in sale probability per percentage point of rate lock-in, estimated on approximately 50 million loans. The provenance of every external benchmark, including the period on which it was estimated, is recorded (§11).

Regression testing and independent review

The automated regression suite currently comprises approximately 1,000 tests. Each test written to guard a specific defect is demonstrated to fail against the defective code before the fix is accepted, so the suite's protective value is verified rather than presumed. In August 2026 the codebase underwent a structured internal review: ten independent area-by-area examinations produced 48 candidate findings; each was then independently re-examined with runtime reproduction, confirming 45 and refuting 3; all 45 confirmed findings were resolved, with the 17 that affected accepted model output corrected through the quantified-impact review process of §12, one reversible change at a time. The complete findings ledger, including reproduction evidence for each item, is retained as part of the model documentation.

Out-of-sample protocol

A documented out-of-sample protocol specifies two temporal splits (estimation through 2018 with evaluation on 2019–2021, and estimation through 2021 with evaluation on 2022–2025), acceptance criteria (absolute evaluation-period bias within approximately 0.5 CPR points; no incentive cohort deteriorating by more than 2 points), and — deliberately — a register of the protocol's own information-leakage channels, such as external benchmarks published from evaluation-period data. All validation is conditional: realized rates and home prices are supplied, and the behavioral response is what is scored. Execution of the full protocol is a planned exercise; its design and leakage register form part of this documentation set.

11Limitations ↑ top

The suite's documentation practice is to state residual weaknesses explicitly rather than absorb them into the fit. The material limitations follow.

12Governance and change control ↑ top

This paper describes the Octant Prepay Model (OPM), the modeling suite behind the forecaster at octantanalytics.com. It is model documentation, not investment advice; reported figures describe historical fit and validation evidence, not guarantees of future performance. Detailed behavioral documentation, policy catalogs, the review ledger, and the out-of-sample protocol are maintained alongside the model source and are available to validation teams on request.