Landmark Analysis in Clinical Research: Dynamic Prediction and Time-Varying Risk Sets
Landmark analysis estimates future risk among patients who are event-free and observable at a prespecified landmark time, using information available up to that time. Repeating the analysis across landmark times supports dynamic prediction as patient histories evolve. The design clarifies the risk set and avoids immortal-time errors, but it requires explicit landmark eligibility, prediction horizon, handling of time-varying measurements, missing data, and landmark-specific calibration.
Clinical risk is not fixed at baseline. A patient’s laboratory values, treatment response, imaging findings, hospitalization history, and functional status can change during follow-up. A prediction made at diagnosis therefore answers a different question from a prediction made six months later after the patient has accumulated new information. Landmark analysis provides a transparent way to formulate the latter question: among patients who have reached a defined time without the event and with sufficient observation, what is their risk during a subsequent prediction window?
The method is useful when the analysis must use information available at a clinical decision point. It can be applied to dynamic prognostic research, survivorship studies, transplant follow-up, oncology response assessment, chronic disease monitoring, and repeated clinical prediction. It is not a replacement for every time-dependent survival model. Its strength is the clarity of the target population and prediction time; its limitations arise when landmark spacing, measurement timing, or selection into the landmark risk set is poorly handled.
What is a landmark time?
A landmark time is a prespecified time point at which eligibility is assessed and predictors are assembled. For example, an analysis may define a 90-day landmark and estimate the probability of hospitalization between day 90 and day 365 among patients alive, event-free, and under observation at day 90. The outcome clock for that prediction window begins at the landmark, not at study entry.
The landmark definition must be clinically meaningful and operationally reproducible. It should specify whether a patient must be alive, free of the event, enrolled, untreated, or measured within a permitted window. If a biomarker is required, the protocol should state how much delay is allowed and what happens when the measurement is missing. These decisions define the target population and cannot be left to the software default.
Landmark analysis differs from simply adding a time-updated covariate to a Cox model. A time-dependent Cox model estimates an association over continuously changing risk sets under its own assumptions. A landmark analysis creates a sequence of conditional prediction problems, each beginning among subjects who meet the landmark criteria. The two approaches may answer related but not identical questions and should not be described as interchangeable.
Dynamic prediction versus baseline prediction
Baseline prediction uses variables measured at or before time zero to estimate a future outcome. Dynamic prediction updates the information set at later times. A dynamic prediction statement must contain at least four elements: the landmark time, the eligible population, the prediction horizon, and the information available at the landmark.
For example, “the 12-month risk after a 6-month landmark among patients who remain event-free and have a recorded assessment at month 6” is more precise than “the model predicts one-year risk.” The former identifies when prediction occurs and who is included. If the horizon is fixed at 12 months after every landmark, the clinical question is different from a model that predicts risk by a common calendar date.
Dynamic prediction can be reported as a sequence of risk estimates, a model that is refit or updated at each landmark, or a joint model that links longitudinal measurements to an event process. The choice should follow the data structure. A landmark approach is often attractive because it is modular and interpretable, but it may discard information between landmarks or produce unstable estimates when few patients remain at later times.
Defining the landmark risk set
The risk set is the most important design object in a landmark analysis. At landmark time s, include only subjects who satisfy the prespecified eligibility rules at s. Patients who experienced the event before s are not included in a prediction of a future event among event-free survivors. Patients lost before s may also be excluded, but the exclusion process should be described because it can change the target population.
Eligibility is not the same as causal treatment assignment. If the landmark analysis compares treatment strategies, conditioning on survival or being event-free can create selection issues, especially when treatment affects the probability of reaching the landmark. The analysis should then be framed as a conditional prognostic or descriptive prediction question unless additional causal assumptions and methods are justified.
Patients can contribute to more than one landmark if the design allows repeated landmark analyses. This creates dependence between estimates from different landmark datasets. Confidence intervals and model comparisons should account for the repeated use of individuals when a joint summary or formal comparison is reported. A simple plot of independent-looking point estimates can otherwise overstate the information available.
Time-varying measurements and measurement windows
Clinical variables are rarely measured at exactly the same time for every patient. A landmark protocol should define an observation window, such as measurements obtained shortly before or after the landmark, and distinguish values available before the prediction decision from values observed afterward. Using information that became available after the stated prediction time creates temporal leakage.
Last-observation-carried-forward is not a universal solution to irregular measurements. A carried-forward value may be clinically implausible when a marker changes rapidly. Complete-case analysis can also select patients who are monitored more intensively or are healthier enough to return for assessment. Multiple imputation, joint longitudinal-survival models, inverse-probability approaches, or sensitivity analyses may be appropriate depending on the missingness mechanism and analysis objective.
When a biomarker is measured repeatedly, researchers should specify whether the model uses the latest value, a slope, a cumulative average, a change from baseline, or a latent trajectory. These choices correspond to different clinical questions. A dynamic prediction model should not present a generic “time-varying covariate” label without explaining the feature construction and its timing.
Prediction horizons and competing events
The prediction horizon should be defined relative to the landmark. Short horizons may support immediate clinical decisions, whereas longer horizons can be useful for survivorship planning but are more sensitive to censoring and changes in care. The number of events available within each landmark-specific horizon affects precision, calibration assessment, and model complexity.
Competing events require an explicit estimand. If death prevents hospitalization from being observed, a cause-specific survival probability and a cumulative incidence probability answer different questions. A landmark prediction can incorporate competing events through an appropriate multi-state or competing-risks framework, but the event definition and interpretation must remain consistent across landmarks. A model that censors competing events without explaining the estimand can be clinically misleading.
Administrative censoring and loss to follow-up should be handled using methods compatible with the prediction target. If censoring is informative, ordinary Kaplan–Meier or unadjusted landmark survival estimates may be biased. Report the follow-up distribution, censoring rules, and whether inverse-probability-of-censoring weighting or another adjustment was used.
Modeling and validating landmark predictions
A landmark model may use Cox regression, flexible parametric survival regression, logistic regression for a fixed horizon, competing-risks regression, or another model suited to the outcome and horizon. The model family should be chosen after defining the estimand. A Cox model estimates relative hazard under proportionality assumptions; a fixed-horizon binary model estimates a risk contrast or probability at a specified time. These are not the same output.
Validation should be performed within the same landmark framework as the intended use. Discrimination can be summarized at the relevant horizon, but calibration is essential: compare predicted and observed risk among patients in the landmark risk set. Calibration-in-the-large, calibration slope, flexible calibration curves, Brier score, and time-dependent performance measures may be useful. A model that performs well at baseline may be poorly calibrated at a later landmark because the case mix has changed.
Internal validation should preserve the landmark construction. If bootstrapping or cross-validation is used, the full process of selecting eligible subjects, constructing time-varying predictors, fitting the model, and calculating predictions should occur inside the resampling loop when appropriate. Evaluating a model after the landmark dataset has already been selected can underestimate optimism if selection and feature construction are ignored.
Evidence summary table
| Design issue | Recommended practice | Interpretation boundary |
|---|---|---|
| Landmark time | Prespecify the time point or grid and define the prediction horizon relative to it. | Changing landmarks after seeing performance creates a different analysis and can inflate optimism. |
| Risk set | Include only patients meeting the event-free and observation criteria at the landmark. | Conditioning on landmark survival changes the target population; it is not automatically a causal comparison. |
| Information timing | Use only measurements available by the prediction decision time. | Post-landmark measurements create temporal leakage and overly optimistic prediction. |
| Prediction horizon | State whether the horizon is fixed after each landmark or tied to a calendar date. | Different horizons imply different clinical questions and event rates. |
| Competing events | Define whether the target is cause-specific survival, cumulative incidence, or a multi-state probability. | Censoring a competing event without an estimand does not produce a universally interpretable risk. |
| Missing measurements | Describe the measurement window and use a missing-data strategy compatible with the target population. | Patients with observed follow-up measurements may differ systematically from those without them. |
| Validation | Assess landmark-specific calibration, discrimination, and optimism using the full prediction pipeline. | Baseline performance cannot be assumed to represent later dynamic performance. |
| Repeated landmarks | Account for reuse of participants when summarizing estimates across landmark times. | Point estimates from separate landmarks are not necessarily independent. |
Actionable Steps: Build a landmark prediction analysis
| Step | Action | Quality gate |
|---|---|---|
| Step 1 | Write the estimand as: among whom, at which landmark, predicting what outcome, over which horizon? | The eligible population, event definition, horizon, and prediction time are explicit. |
| Step 2 | Define the landmark risk set and construct a data-availability rule for every predictor. | No variable uses information that became available after the prediction decision. |
| Step 3 | Choose the outcome model and competing-event strategy before examining landmark-specific performance. | The model output matches the clinical estimand rather than a convenient software default. |
| Step 4 | Fit and validate predictions at the intended landmark times, including calibration and optimism assessment. | Resampling repeats eligibility, feature construction, model fitting, and prediction when needed. |
| Step 5 | Report performance by landmark, missingness, censoring, competing events, and clinically relevant subgroups. | Readers can judge transportability and identify where dynamic predictions may fail. |
Common failure modes
The first failure is treating a landmark analysis as a simple subset analysis. A subset created at a later time has a different risk set and potentially different selection mechanisms. The manuscript should explain why patients enter the analysis and what the conditional prediction statement means.
The second failure is using future information. A laboratory value recorded after the landmark may be strongly predictive, but it cannot be used for a model intended to support a decision at the landmark. Timestamp checks and a prespecified feature-freeze rule are practical safeguards.
The third failure is reporting only a discrimination statistic. Dynamic prediction can rank patients reasonably while systematically overestimating or underestimating absolute risk. Landmark-specific calibration and observed-to-expected summaries are necessary when the model is intended to guide decisions or follow-up intensity.
The fourth failure is ignoring competing events. A model may appear to predict hospitalization well until death prevents hospitalization from occurring. The target estimand should state how death is treated, and the prediction output should match that choice.
The fifth failure is fitting separate models at many landmarks without considering multiplicity, sparse late risk sets, or dependence across estimates. A smaller number of clinically justified landmarks, shrinkage, joint modeling, or prespecified summaries may be preferable to a dense grid of unstable analyses.
Reporting and workflow considerations
A reproducible report should define time zero, each landmark, the eligibility rule, the prediction horizon, predictor measurement windows, feature construction, censoring, competing events, missing data, and model specification. It should show how many patients were eligible at every landmark and how many events occurred during each prediction window. These counts determine whether late-landmark performance estimates are credible.
For prediction models, follow TRIPOD-oriented reporting principles and provide calibration alongside discrimination. For observational prognostic research, distinguish conditional prognosis from a causal treatment effect. If treatment changes over time, a landmark prediction may be clinically useful without identifying the effect of a treatment strategy. Causal questions may require a target-trial, marginal structural, or other time-varying causal framework.
Researcher's Toolkit: Strengthen Dynamic Prediction Research
Lingcore SCI supports the evidence and reporting workflow around landmark analyses:
- Paper Analyzer: Extract landmark definitions, risk-set rules, prediction horizons, competing-event handling, and calibration results from published studies.
- Review Builder: Organize a citation-linked methods review comparing landmarking, joint models, time-dependent survival models, and multi-state prediction.
- Journal Matcher: Identify journals suited to clinical prediction, survival analysis, epidemiology, outcomes research, and biostatistical methods.
These tools can improve organization and reporting consistency, but investigators remain responsible for timestamps, data quality, assumptions, model validation, and scientific interpretation.
Conclusion
Landmark analysis turns a changing clinical history into a sequence of explicit prediction problems. Its value comes from defining who is eligible at each decision time, freezing the information set, aligning the horizon with the clinical question, and validating absolute risk in the appropriate landmark population. Used carefully, it can make dynamic prediction more transparent and clinically interpretable. Used casually, it can introduce selection bias, temporal leakage, unstable late estimates, or ambiguous risk statements. A rigorous landmark report makes each of these choices visible.
LINGCORE SCI