International Institute for Musculoskeletal Health Education

IIMHE response to the updated MelioGuide article (12 August 2026) regarding REMS scanning

A detailed scientific and methodological response to ‘REMS Bone Scan (Echolight): Is This the DEXA Alternative?’

International Institute for Musculoskeletal Health Education · Publication edition · 16 August 2026

Executive summary

IIMHE’s central conclusion is deliberately balanced: demographic inputs materially influence REMS outputs, but the available evidence does not establish that REMS is only a demographic calculator, nor does it quantify the skeletal contribution with sufficient certainty. [1–12]

REMS and DXA are not interchangeable; individual decisions and monitoring require a consistent modality, protocol and qualified interpretation. [4,8,9]

Chan, Bobelyak and Baquir provide convergent evidence of strong demographic association. Their R² values describe variance in study populations, not a percentage of an individual’s result. [4,11,12]

Birch reports substantial variation within age/BMI segments, while Conversano reports a lower demographic contribution in 16,000 scans. Neither residual variance nor within-segment spread can automatically be labelled skeletal information. [2,10]

Montatore shows group-level discrimination after matching mean age and BMI, but its cross-sectional design, lack of baseline and absence of DXA prevent causal, accuracy or monitoring claims. [5]

Independent access to pre-integration ultrasound-derived data and the proprietary weighting architecture is the most important unresolved requirement. [3,4,8,10,11]

Public discussion should distinguish routine DXA acquisition from VFA, and REMS B-mode localisation from diagnostic vertebral imaging. [1,8]

Contents

  • Executive summary
  • 1. Changes in the MelioGuide article since May 2026
  • 2. New and persistent issues in the updated MelioGuide article
  • 3. The Baquir and Montatore studies: methodological analysis
  • 4. Continuing scientific questions
  • 5. Ongoing research and its evidential status
  • 6. Points on which IIMHE agrees with MelioGuide
  • 7. Wider considerations for an evidence-led debate
  • 8. Methodological framework for interpreting the evidence
  • 9. An independent research programme to resolve the dispute
  • 10. Specific amendments requested from MelioGuide
  • Conclusion
  • Declarations
  • References

Introduction

The International Institute for Musculoskeletal Health Education (IIMHE) welcomes the continued engagement of Margaret Martin and the MelioGuide platform with the evolving literature on radiofrequency echographic multi-spectrometry (REMS). The article updated on 12 August 2026 makes a substantive effort to incorporate new evidence and address criticisms raised after its original publication, including points in IIMHE’s May 2026 response. We acknowledge those revisions and the editorial willingness they demonstrate. [1]

Since IIMHE published its first response on 1 May 2026, the debate has advanced. Birch, Tognarini, McCloud and Button published a letter in Osteoporosis International on 11 August 2026, accompanied by a rebuttal from Pocock, Bobelyak, Vaculik and Stepan. Baquir and colleagues published an independent Australian validation study in Bone, and Montatore and colleagues published a comparative study of men receiving androgen deprivation therapy in Aging Clinical and Experimental Research. [2–5] The updated MelioGuide article addresses the Birch and Pocock correspondence but does not yet analyse the Baquir or Montatore studies. This response therefore considers the complete evidence available to 15 August 2026. [1–5]

IIMHE’s position remains evidence-led and technology-neutral. The central question is how REMS can be used safely and appropriately in fracture-risk assessment. That question must be answered through transparent methods, reproducible research and clinically responsible interpretation, with the same critical standards applied to evidence both for and against REMS. [1–5,8–12]

1. Changes in the MelioGuide article since May 2026

1.1 Corrections accepted following IIMHE’s May 2026 response

The original article stated that there was no published evidence that REMS could track treatment response. That statement has been removed. The updated article now describes the Semeraro romosozumab cohort and the Caffarelli denosumab/aromatase-inhibitor cohort, with appropriate recognition that both are observational studies requiring independent replication. These are material additions. [6,7]

The article also acknowledges REMS in Italian inter-society fragility-fracture guidance and the Polish Society for Rheumatology’s 2026 recommendations, and it accurately explains the limited regulatory meaning of the US CPT Category III code 0815T. These changes improve the description of REMS’s current status without overstating it. [1]

The Zambito REMS practice-parameters paper is now considered more substantively. The article identifies circumstances in which REMS may have defined utility, including situations where DXA is impractical or its interpretation is compromised by degenerative change, spinal hardware or other artefacts; where ionising radiation should be minimised; and in selected peri-operative and chronic-kidney-disease pathways. These are use cases emphasised by IIMHE, while the practice parameters also make clear that REMS has limitations and does not replace vertebral fracture assessment. [8]

The 2026 systematic review and meta-analysis by Liu and Liu, which pooled 17 studies and more than 8,700 participants, is now cited and discussed. It found pooled REMS–DXA correlations of 0.82 at the femoral neck and 0.77 at the lumbar spine, while also reporting very high heterogeneity. Correlation supports convergent validity at population level; it does not establish interchangeability for individual patients. [9]

The methodological clarification by Conversano, Pisani and Casciaro, based on 16,000 REMS scans, is now included. Its estimates place the contribution of demographic and anthropometric variables at approximately 60%, compared with more than 90% in Chan and colleagues. The difference is unresolved and should be examined rather than averaged away. [10,11]

In this updated article, the Trabecular Bone Score (TBS) is now subjected to the same level of critical scrutiny applied to REMS. The article notes that TBS is a texture measurement from a two-dimensional DEXA image rather than a direct visualisation of trabeculae, that it is affected by body size, soft tissue and image noise, that its evidence base is weakest in men, and that ISCD advises against using it in isolation or for routine longitudinal monitoring. This even-handedness was absent from the original article, and we welcome its inclusion. [1]

1.2 The Birch and Pocock correspondence in Osteoporosis International

The updated article cites the Birch letter and the Pocock rebuttal. [2,3] The Birch letter presents large real-world datasets showing substantial BMD and T-score variability within age/BMI segments and proposes that analyses intended to characterise the full within-segment range may require samples approaching the size of the REMS reference segments. The authors explicitly describe this as a postulate and say the observed trend suggests that close to 100 observations may be needed. It is an important, testable methodological argument, but it is not an established universal sample-size rule for every regression question. [2,3]

1.3 The conflict-of-interest description remains inaccurate

The article continues to describe IIMHE as a UK organisation partnered with a REMS distributor. IIMHE has a relationship with OsteoScan UK Ltd, an independent bone-health assessment provider. IIMHE states that it has never been partnered with a UK or overseas distributor of REMS equipment and has no commercial relationship with Echolight or a REMS distributor. The distinction is relevant to readers assessing institutional interests and should be corrected. IIMHE asks MelioGuide to describe the relationship accurately and to link to IIMHE’s own disclosure statement. [1]

2. New and persistent issues in the updated MelioGuide article

2.1 The Pocock rebuttal requires more precise representation

The updated article discusses the Pocock, Bobelyak, Vaculik and Stepan rebuttal alongside the Chan paper in a way that may imply substantial authorial overlap. Only Pocock is common to both author groups. More importantly, the rebuttal accepts that entered demographics materially influence REMS-BMD and Fragility Score outputs; its dispute concerns what can legitimately be inferred about the residual variation and the relative weighting of demographic and ultrasound-derived components. [3,11]

The article notes that the rebuttal declares no conflicts of interest, while Pocock disclosed ownership of a DXA service in the Chan paper. The appropriate point is one of consistent disclosure across publications and commentary, not proof that the scientific analysis is biased. Commercial involvement can inform how evidence is weighed, but it neither invalidates nor validates a result. IIMHE applies the same principle to REMS-associated and DXA-associated authors, including its own contributors. [3,11]

The rebuttal makes three legitimate points: residual variance not explained by demographics cannot automatically be labelled skeletal information; deliberate demographic manipulation does not by itself quantify the weighting of each algorithmic component; and independent access to raw spectral data and the underlying algorithms is required to determine those contributions. [3] These points deserve direct engagement, even where IIMHE disagrees with the broader interpretation. [3,11]

2.2 The Birch sample-size argument should be considered, but not overstated

The Birch letter raises a methodological challenge to the Bobelyak and Chan datasets: their participants are distributed across multiple age/BMI segments, leaving few observations within many individual cells. [2,11,12] This matters if the research objective is to estimate the full within-segment distribution or reliably observe its extremes. It does not follow, however, that every regression of a continuous outcome on age and weight is invalid unless every algorithmic segment contains 100 participants. The estimand, model form, precision and external validation all matter. [2,11,12]

Birch and colleagues analysed 5,137 female patients from established REMS services, comprising 6,389 lumbar-spine and 8,744 hip scans. In segments with more than 50 scans, T-score ranges exceeded 2.0 T-score units; many segments spanned normal, osteopenic and osteoporotic thresholds. The observation demonstrates substantial within-segment variation, contrary to a purely deterministic demographic-output model. It does not, by itself, establish that the variation is skeletal rather than a combination of skeletal signal, acquisition, operator and measurement noise. The correct next step is independent replication with prespecified within-segment analyses. [2]

2.3 The left–right hip correlation requires cautious interpretation

The article describes the Chan left–right femoral-neck correlation of r = 0.99 as biologically implausible and contrasts it with DXA correlations around 0.90–0.93. [1,11] The very high REMS correlation is a valid signal that merits investigation, particularly because identical demographic inputs are used for both hips. Correlation alone cannot partition the effect into demographic compression, bilateral biology, acquisition or algorithmic processing. A designed same-patient study that models these sources separately would be more informative than a categorical inference from r alone. [1,11]

2.4 The FRAX double-counting concern is primarily a misuse warning

Entering a REMS-derived femoral-neck BMD into FRAX could theoretically duplicate or confound demographic effects because age and weight already contribute to FRAX. However, standard REMS reports provide a separate five-year fracture-risk estimate rather than a FRAX probability. The practical conclusion should therefore be explicit: REMS-BMD should not be entered into FRAX unless and until that use is validated. The concern is important as a warning against off-label calculation, not evidence of an error occurring in standard REMS reporting. [1,11]

2.5 Fragility Score construction and clinical prediction are separate questions

Chan and colleagues found that a quadratic age function predicted more than 95% of femoral-neck Fragility Score variance and more than 80% at the lumbar spine. [11] Pisani and colleagues reported five-year fracture-prediction AUCs of approximately 0.78–0.81. [13] A strongly age-structured score can still predict fractures because age is itself a major fracture-risk determinant. Predictive discrimination does not prove that the score adds value beyond age, but demographic dependence does not by itself prove that it adds none. [11,13,28]

The necessary analysis is incremental: does Fragility Score improve prediction after prespecified adjustment for age, BMI and established clinical risk factors, and is that improvement replicated independently with calibration as well as discrimination assessed? The Pisani paper reports retention of predictive value after adjustment for age and BMI, but independent external validation remains necessary. IIMHE therefore supports neither dismissal nor uncritical adoption of the score. [11,13,28]

2.6 REMS preparation is described too broadly

The article presents fasting for several hours as a general preparation requirement. Fasting relates principally to lumbar-spine acquisition, where bowel gas can compromise the transabdominal ultrasound window; preparation varies with the patient, protocol and operator. The wording should distinguish lumbar-spine preparation from hip acquisition and direct patients to the instructions of the service performing the scan. [1,8]

3. The Baquir and Montatore studies: methodological analysis

Baquir and colleagues’ Australian validation study was accepted on 9 July 2026 and published online in Bone in August 2026. [4] It is an important independent contribution because it examines agreement, diagnostic classification and demographic determinants in a real-world hospital cohort. The updated MelioGuide article does not yet discuss it. [4]

Montatore and colleagues’ study was published online on 4 August 2026 in Aging Clinical and Experimental Research. It compares men with prostate cancer receiving androgen deprivation therapy (ADT) with age- and BMI-matched healthy controls. [5] It is also absent from the current MelioGuide analysis. [5]

Scientific balance requires both studies to be considered on their own designs and limitations. Neither should be treated as a decisive answer to the demographic-dependence question. [4,5]

3.1 Baquir et al.

3.1.1 Study design and population

Baquir and colleagues analysed 69 participants (27 men and 42 women) from Westmead Hospital who had lumbar-spine and proximal-femur REMS and DXA assessments. The cohort had a mean age of 61.6 years, mean weight of 73.3 kg and mean BMI of 27.2 kg/m². It included 37 White participants, 24 Asian participants and eight participants from other groups; 59.4% had chronic inflammatory disease, 55.1% had a prior fragility fracture and 37.6% were receiving osteoporosis medication. DXA was performed on a GE Lunar iDXA or a Hologic Horizon A. [4]

The reported femoral-neck T-scores range from −4.0 to +3.3, a span of 7.3 T-score units, and lumbar-spine T-scores range from −3.8 to +5.6, a span of 9.4 units. These are very wide ranges that warrant inspection for machine, population, positioning, degenerative and analytical effects. [4]

3.1.2 Central finding and methodological context

A regression model containing age and weight explained approximately 90–91% of REMS femoral-neck BMD variance and 88% of lumbar-spine variance, compared with approximately 33–35% for DXA at corresponding sites. [4] This independently reproduces the direction of the Bobelyak and Chan findings. The result is important evidence that demographic inputs are strongly associated with reported REMS-BMD in this dataset. [4,10–12]

It does not mean that 90% of any individual’s reported value was generated by demographics, nor does the unexplained fraction automatically represent bone. R² is a population-level measure of explained outcome variance under a specified model. Interpretation depends on sampling, model stability, measurement error, confounding and external validation. [4,10–12]

3.1.3 Sample size and segment representation

The 69 participants span broad age and BMI ranges and therefore populate many of the 42 age/BMI segments described for the REMS reference architecture. Individual cells are necessarily sparse. This limits the study’s ability to characterise within-segment ranges, assess outliers or examine effect modification across demographic groups. The Birch dataset suggests that observed ranges widen as segment samples increase, but the proposed threshold of approximately 100 observations remains a hypothesis for testing rather than a general requirement that invalidates the regression. [2,4,10]

3.1.4 Two-machine DXA limitation

The study used a GE Lunar iDXA and a Hologic Horizon A without cross-calibration. The authors included scanner type in regression models, where it was associated with 20% of femoral-neck and 16% of lumbar-spine DXA variance. [4] Because machine was also linked to department and therefore potentially to patient selection, site, operator and protocol, this association should not be interpreted as a pure manufacturer effect. [4,14,15]

Inter-system DXA differences are well established. Universal standardisation reduced, but did not eliminate, differences between systems, particularly at the spine. [14,15] The resulting densitometry framework requires appropriate cross-calibration before quantitative inter-system comparison. The absence of cross-calibration is a material limitation when REMS is compared with a mixed DXA reference, but it does not invalidate the finding that REMS and DXA were not interchangeable in this cohort. [4,14,15]

The scanner variable must not be arithmetically added to the demographic R² and subtracted from 100% to estimate a residual skeletal component. Predictor contributions can overlap, and scanner may proxy site or population differences. The defensible conclusion is narrower: the DXA comparator contains important heterogeneity that widens uncertainty around direct modality comparisons. [4]

The comparison is therefore not a pure contest between ultrasound signal and X-ray physics. Both reported outcomes contain biological, technical and analytical sources of variation; the proportions cannot be partitioned from this study alone. [4,14,15]

3.1.5 LASSO models and optimism correction

The LASSO-regularised models for REMS retained high corrected R² values (approximately 92% at the femoral neck and 87% at the lumbar spine), whereas corrected DXA performance fell to approximately 12% and 27%. [4] The fall in DXA performance indicates model instability and poor expected out-of-sample prediction in this small dataset. It does not support an inference that the true demographic contribution to DXA must be higher than the uncorrected estimate. The direction and magnitude remain uncertain until tested in a larger independent cohort. [4]

3.1.6 Population diversity, reference data and measured BMD

The cohort is too small to examine performance across ethnic groups or test whether relationships differ by ancestry. Reference populations affect T- and Z-score derivation; they do not directly create disagreement in measured BMD expressed in g/cm². The wide Bland–Altman limits for BMD are more appropriately discussed in relation to calibration, scanner, acquisition, site, operator, patient mix and modality. Reference-database mismatch should be analysed separately when diagnostic categories or standardised scores are compared. [4,14,15]

3.1.7 Medication and clinical heterogeneity

More than one-third of participants were receiving active osteoporosis medication. Treatment status may represent genuine skeletal variation not captured by age and weight, but the study was not powered to determine whether the REMS residual tracked treatment exposure. Medication, inflammatory disease and prior fracture should be prespecified covariates or stratification variables in larger validation datasets. Their omission limits interpretation; it does not demonstrate that treatment-related signal is absent or present. [4]

3.1.8 Operator training and acquisition context

The paper reports that the two REMS operators received specific manufacturer training and had at least three months of experience before scanning. [4] This is relevant and should be represented accurately. The publication does not provide a detailed competency log or independent acquisition-quality audit, so operator effects cannot be fully quantified; equally, the available description does not justify characterising the operators as untrained. [4]

3.1.9 What Baquir establishes — and what it does not

Baquir demonstrates that age and weight explain a high proportion of REMS-BMD variance in this small hospital cohort; that REMS and DXA showed wide individual-level limits of agreement; and that the modalities should not be used interchangeably. It supports calls for manufacturer clarification, external validation and access to pre-integration spectral information. [4]

It does not establish that REMS measures nothing beyond a demographic prediction. The sample is too small to estimate the full range within demographic strata, investigate subgroup performance or isolate the effects of two uncross-calibrated DXA systems and clinical heterogeneity. The result should be taken seriously without being made to answer a question its design cannot resolve. [4]

3.2 Montatore et al.

3.2.1 Study design and population

Montatore and colleagues report a single-centre comparative observational study enrolling 143 White men: 83 with histologically confirmed prostate cancer receiving ADT for at least six months and 60 healthy controls matched on mean age and BMI. REMS was performed at the lumbar spine and femoral neck. The analysis is cross-sectional: there were no pre-ADT baseline measurements and no comparative DXA scans. [5]

Chan, Bobelyak and Baquir estimate associations between REMS outputs and demographic variables. That approach quantifies explained variance under the fitted model but cannot determine whether residual variance is skeletal signal. Montatore asks a different question: can REMS distinguish groups expected, on prior evidence, to differ in bone status despite similar mean age and BMI? [4,5,10–12]

The known-groups logic is useful, but it should be stated cautiously:

  • ADT is associated with bone loss and increased fracture risk; this is established independently of REMS. [5]
  • The study groups have similar mean age and BMI, reducing, but not eliminating, demographic confounding. [5]
  • If reported REMS-BMD were solely and deterministically generated from those variables, little group separation would be expected. [5]
  • Observed separation is therefore evidence against a solely deterministic age/BMI output at group level. It is not proof of measurement accuracy, causal ADT effects or longitudinal sensitivity. [5]

3.2.2 Results

PCa on ADT (n = 83)Control (n = 60)Difference
Age, y74.2 ± 7.073.5 ± 7.30.7
BMI, kg/m²25.3 ± 2.925.3 ± 3.10.0
LS BMD, g/cm²0.885 ± 0.080.945 ± 0.110.060 (6.35%)
FN BMD, g/cm²0.691 ± 0.080.715 ± 0.100.024 (3.36%)
LS T-score−1.9 ± 0.8−1.3 ± 1.00.6
FN T-score−1.8 ± 0.7−1.5 ± 0.70.3
LS Z-score−0.7 ± 0.7−0.2 ± 0.90.5
FN Z-score−0.7 ± 0.6−0.2 ± 0.90.5

PCa on ADT = participants with prostate cancer receiving androgen deprivation therapy. Values are mean ± SD. Source: Montatore et al. [5]

The groups differed by only 0.7 years in mean age and not at all in mean BMI, while mean REMS-BMD was 6.35% lower at the lumbar spine and 3.36% lower at the femoral neck in the ADT group. These differences are compatible with known-groups discrimination. They cannot be attributed causally to ADT because baseline status and residual clinical confounders were not measured. [5]

Z-scores are relevant because they express deviation from an age-matched reference. A purely deterministic age/BMI algorithm would be expected to compress group differences after matching; however, real Z-scores have a distribution and need not cluster at zero even when demographics contribute strongly. The observed difference therefore argues against a solely demographic output at group level, not against a substantial demographic contribution. [5,14–16]

Controls had mean Z-scores of −0.2 at both sites and the ADT group −0.7. The consistency across two anatomical sites is noteworthy, but it remains an association requiring confirmation with baseline and comparator measurements. [5]

The Birch within-segment sample-size argument is not directly transferable to this comparison of group means. Larger samples improve precision in both settings, but estimating a distribution’s extremes is a different task from testing a prespecified mean difference. The present group sizes may detect a difference; they do not eliminate confounding or establish instrument accuracy. [2,5]

3.2.3 Study limitations

The study reports BMI rather than separate weight and height matching. Mean BMI equivalence does not ensure identical weight distributions, which matters because weight is an entered REMS variable. The groups may also differ in cancer-related health, medication, activity, comorbidity and selection factors. [5]

There is no DXA comparison and no pre-ADT baseline. The study therefore establishes cross-sectional discrimination, not agreement, causal change, longitudinal monitoring performance or accuracy. [5]

Table 3 appears to report a femoral-neck T-score difference of 0.49, whereas the group means in Table 1 imply approximately 0.3. This should be described as an apparent internal inconsistency, likely typographical in the article-in-press version. Table 1 is internally consistent with the accompanying group means, but its external accuracy cannot be established from the manuscript alone. IIMHE has notified the authors. [5]

Subject to those limitations, the design adds a useful known-groups perspective to a literature dominated by cross-sectional demographic regression. It identifies a study design worth strengthening with pre-ADT baselines, repeated measures, a calibrated DXA comparator and independent analysis. [5]

3.2.4 What Montatore establishes — and what it does not

Montatore is a useful contribution containing an apparent reporting inconsistency that should be corrected. That inconsistency does not alter the principal group means but should not be minimised or sensationalised. [5]

The study shows a group-level REMS difference not explained by mean age or BMI and consistent with established skeletal effects of ADT. The Z-score separation is informative because it is expressed relative to age-matched expectation. [5]

The study does not prove that REMS accurately measures ADT-related bone loss, does not validate longitudinal monitoring and does not quantify the ultrasound contribution to the algorithm. Its value is to show that demographic regression is not the only informative design. Larger prospective studies with baseline REMS and DXA, prespecified covariates and blinded independent analysis are needed. [5]

4. Continuing scientific questions

4.1 The Chan volunteer-manipulation experiment

In five healthy volunteers, Chan and colleagues altered entered age or weight while the underlying skeleton remained unchanged and observed predictable changes in REMS-BMD. [11] This within-person manipulation robustly demonstrates that demographic inputs can materially change the reported output. The sample of five limits precision and population-level generalisation; it does not erase the experimental observation. [11]

The experiment does not, however, quantify how much of a correctly entered patient’s output is attributable to demographic weighting rather than ultrasound-derived information. It also cannot distinguish every possible implementation mechanism within the proprietary processing pipeline. Those questions require access to pre-integration spectral outputs and a prespecified technical validation design. [3,11]

Reference-system effects have a DXA precedent. Faulkner, Roberts and McClung found that the same women assessed on Lunar and Hologic systems could receive systematically different femoral-neck T-scores because of differences in normative data and scoring models. [16] This illustrates why measurement and scoring layers must be distinguished; it does not make demographic manipulation clinically irrelevant. [16]

The skeleton was unchanged while the reference frame changed. The appropriate response was reference harmonisation and standardisation, not the conclusion that DXA was merely demographic. The analogy supports investigation of REMS’s processing layers, but it cannot establish that REMS and DXA are technically equivalent. [14–16]

For REMS, the essential information remains inaccessible to independent investigators: the raw ultrasound-derived output before integration and the exact weighting of demographic variables. Without it, neither advocates nor critics can definitively partition the reported BMD. [3,10,11]

The Pocock rebuttal is correct that observed variability alone cannot quantify the spectral or demographic weighting. [3] The Chan experiment is correct that entered demographics alter the final value. Both observations can be true, and both strengthen the case for independent technical transparency. [3,11]

4.2 The Conversano analysis deserves careful consideration

Conversano, Pisani and Casciaro analysed 8,000 lumbar-spine and 8,000 proximal-femur scans from White women using separate development and test datasets. [10] Their finding, substantial demographic contribution but typically at least 40% variance not explained by the reported demographic model, is the largest analysis currently available. Its manufacturer affiliation and restricted population require appropriate caution. Equally, its size, held-out testing and direct focus on the disputed question mean it should not be reduced to a passing ‘roughly 60%’ figure. The residual cannot automatically be labelled bone; it may include ultrasound-derived skeletal information, acquisition effects and noise. [10]

4.3 Demographic prediction in DXA is relevant but not exculpatory

Demographic and anthropometric models can explain a meaningful proportion of DXA-BMD variance. That context prevents the mistaken assumption that any association between age, weight and a densitometric output proves the absence of measurement. It does not answer whether REMS relies excessively on those inputs for individual assessment. [4,10,11,14–16]

Variance estimates from different datasets and models must not be added or compared as though they were mutually exclusive components. Machine, site, population and demographic variables can overlap. The Baquir scanner association is therefore a warning about comparator heterogeneity, not a percentage that can be subtracted to reveal ‘true skeleton’. [4]

The appropriate comparison uses large, independently analysed datasets, comparable populations, prespecified models, cross-calibrated DXA systems and validation outside the development cohort. Current REMS studies differ materially in sample size, case mix and analysis. The 60% and greater-than-90% estimates define an unresolved range, not a settled parameter. [4,10,11]

4.4 DXA history and the development of measurement standards

DXA did not enter clinical use in the late 1980s as a fully standardised, artefact-free or universally interpretable technology. Its limitations were identified through decades of method development. That history is relevant because it shows how a field should respond to uncertainty: characterise error, standardise acquisition, cross-calibrate systems, validate reference data and define precision requirements. [14–25]

YearFindingWhat the field did
1992Areal BMD is confounded by bone size; a larger bone reads denser at constant volumetric density (Carter [18])Limitation characterised and incorporated into interpretation; size-correction methods developed for paediatrics
1992Non-uniform fat distribution over the vertebrae introduces error into lumbar DXA (Tothill & Pye [19])Error source quantified; soft-tissue effects became a recognised accuracy term
1994Change in body weight and soft-tissue thickness alters measured spine BMD (Tothill & Avenell [20])Weight change recognised as a confounder in longitudinal interpretation
1994Hologic, Lunar and Norland report systematically different BMD for the same patients (Genant [14])International DXA Standardization Committee; sBMD conversion equations
1996Manufacturer normative databases yield different T-scores at the femoral neck (Faulkner [16])Move toward a common reference database; NHANES III adopted for the hip
1996BMD predicts fracture at population level but does not identify which individual will fracture (Marshall [21])Conceptual shift toward multifactorial risk; clinical risk factors added, ultimately FRAX
2000–01Osteoporosis redefined in terms of bone strength, of which density is one component (NIH [22])Bone quality established as a distinct research target; TBS, QCT, HR-pQCT follow
2002Serial DXA monitoring in clinical practice described as subject to recent controversies (ISCD [23–25])Precision assessment mandated; least significant change formalised as a per-centre requirement

Commercial DXA systems were introduced in the late 1980s. By 1992, projected areal BMD was shown to depend on bone size, and non-uniform fat distribution was shown to introduce lumbar measurement error. Weight change and soft-tissue thickness were subsequently identified as longitudinal confounders. [17–20]

In 1994, an international standardisation study quantified systematic differences between manufacturers and developed standardised BMD equations. In 1996, normative databases were shown to yield systematic femoral-neck T-score differences, while a landmark meta-analysis confirmed that BMD predicts fracture risk at population level without identifying exactly which individual will fracture. [14,16,21]

The NIH later defined osteoporosis in terms of compromised bone strength, with density as one component. In 2002, ISCD convened a panel because serial DXA monitoring remained controversial and required formal precision and least-significant-change standards. [22,23]

None of these findings invalidated DXA. Each defined conditions for reliable use, and the field responded with quality assurance. REMS deserves equally rigorous scrutiny. A limitation should neither be concealed nor automatically converted into a verdict that the modality contains no clinically useful skeletal information. [14–25]

4.5 Monitoring and the DXA precedent

The reliability of REMS for longitudinal monitoring remains an open question. The relevant DXA history shows that good monitoring practice depends on more than nominal device precision. [6–8,23–25]

Fifteen years after DXA’s commercial introduction, the ISCD panel concluded that monitoring requires centre-specific in-vivo precision and that an observed change must exceed the least significant change before it is interpreted as biological. [23] Later positions reinforced precision assessment, cross-calibration and consistency of facility, system and protocol. [24,25]

The same discipline should apply to REMS: documented precision, an established least significant change, consistent equipment and acquisition protocol, trained operators and no interpretation of sub-threshold change as biological. The treatment-monitoring studies are encouraging signals, not a substitute for independent prospective validation. [6–8]

4.6 Vertebral imaging, VFA and REMS B-mode

The MelioGuide article is right that REMS does not provide vertebral fracture assessment. It is also important not to equate the low-resolution anteroposterior image used for routine lumbar DXA acquisition with a diagnostic lateral VFA examination. VFA is an additional lateral assessment intended to evaluate vertebral morphology. [1,8]

REMS uses a B-mode image for localisation and acquisition guidance, but current practice parameters state that vertebral structural changes cannot be directly visualised through REMS B-mode. [8] It should therefore not be described as equivalent to VFA, radiography or diagnostic vertebral imaging. Suspected vertebral fracture requires an appropriate imaging pathway. [8]

REMS may be less susceptible than lumbar DXA to some artefacts associated with osteoarthritis, vertebral collapse or extra-skeletal calcification, and small clinical studies support further investigation of that advantage. [26,27] The evidence does not justify a categorical claim that fractured tissue is always excluded or that REMS will necessarily produce a lower BMD. REMS may identify a different value in an artefact-affected spine, but it cannot identify or characterise the fracture itself. [8,26,27]

5. Ongoing research and its evidential status

Additional REMS datasets and manuscripts are in preparation, including larger within-segment analyses and work in complex clinical populations. Until peer reviewed and publicly available, such material should be labelled unpublished and treated as hypothesis-generating. IIMHE does not rely on conference observations or forthcoming manuscripts as proof in this response. It will update its position when methods and results can be independently examined. [2]

6. Points on which IIMHE agrees with MelioGuide

Several parts of the updated article align with IIMHE’s clinical position and should be stated without qualification. [1,4,8]

6.1 REMS and DXA are not interchangeable

REMS and DXA should not be used interchangeably for individual decisions. The available Bland–Altman data, including Baquir, show limits of agreement too wide for substitution. [4,8] A REMS result should not override a DXA-based treatment decision without a specific clinical reason and qualified interpretation. Longitudinal follow-up should normally use the same modality, facility and protocol. Both REMS and DXA require an appropriate professional pathway; an isolated number without clinical context creates risk regardless of technology. [1,4,8]

6.2 A person is more than a T-score

IIMHE agrees that a number is useful only when it informs the right next step. Before booking a REMS assessment, patients should ask who will interpret the findings; whether the report can be shared with their primary clinician or specialist; what pathway follows an abnormal or discordant result; and whether DXA, VFA or other imaging will be recommended when indicated. Those are sensible safeguards and IIMHE endorses them. [8]

7. Wider considerations for an evidence-led debate

The remaining disagreement is not whether REMS should be scrutinised. It should. The question is whether the same evidential standards and careful language are applied to every modality and every research group. [1–12]

7.1 Consistency in evaluating evidence

The article appropriately identifies limitations in REMS research: single-centre designs, small samples, manufacturer-associated authorship, limited external replication and proprietary algorithms. Those factors should reduce certainty. Comparable limitations in studies critical of REMS, including small experimental samples, mixed clinical populations and commercial service interests, should be disclosed and weighed in the same manner. Consistency strengthens the critique; it does not weaken it. [1]

The Chan volunteer experiment contains only five participants, which limits quantitative generalisation, but its within-person manipulation is informative. [11] Conversely, large manufacturer-associated datasets provide greater statistical precision but still require independent replication and access to methods. Commercial involvement is relevant context in both directions and should never substitute for methodological analysis. [2–5,10,11]

7.2 Algorithm transparency is the central unresolved issue

DXA is not perfectly transparent: manufacturers use proprietary hardware and software, and historical reference and calibration differences produced clinically relevant discrepancies. Nevertheless, its measurement principles, reference datasets and standardisation equations have been extensively published and independently tested. [14–16,25]

REMS is less externally auditable. Independent investigators cannot currently determine the weighting of demographic inputs and ultrasound-derived spectral information because the pre-integration output and full algorithms are proprietary. IIMHE regards this as the strongest criticism in the current debate and supports independent access under appropriate data-protection and intellectual-property safeguards. Transparency would allow advocates’ and critics’ competing explanations to be tested rather than inferred. [3,4,8,10,11]

7.3 The claim about demographic outliers remains a hypothesis

MelioGuide argues that a demographically driven system would be least reliable in people whose bones differ most from demographic expectation. [1] That is a reasonable concern, but the premise and magnitude remain disputed. Birch reports substantial T-score variation within individual age/BMI segments, indicating that REMS does not collapse every patient to one segment value. [2] The finding does not prove that all within-segment variation is skeletal. The decisive test is independent, adequately powered outlier validation against appropriate clinical and imaging criteria. [1,2,4,10,11]

7.4 Fragility Score and clinical decision-making

The finding that age predicts much of Fragility Score variance is important and should temper individual interpretation. It does not make an age-structured risk score automatically uninformative. The clinically relevant questions are whether Fragility Score improves prediction beyond age and established risk factors, whether calibration is acceptable and whether any improvement changes management. [11,13]

The broader history supports this distinction. BMD predicts fracture risk incompletely, prompting the incorporation of age, prior fracture, glucocorticoid exposure, smoking and parental hip fracture into tools such as FRAX. [21,28] Demographic structure may be appropriate in a risk model; it is more problematic when its contribution to a quantity presented as BMD cannot be independently quantified. [14–16,21,28]

7.5 The evidence is evolving rather than settled

Current evidence supports three propositions simultaneously: REMS outputs are materially influenced by entered demographics; REMS and DXA are not interchangeable; and available data do not establish that REMS is merely a demographic calculator with no clinically useful skeletal signal. [2–11] A responsible public account should present all three. IIMHE welcomes continued updates by MelioGuide as new peer-reviewed evidence becomes available. [1–11]

8. Methodological framework for interpreting the evidence

The present debate repeatedly uses the same statistical terms to answer different questions. A careful framework is necessary because a study may be valid for one purpose and inadequate for another. The distinctions below are not technical footnotes; they determine what can safely be said to practitioners and patients. [2–5,9–12]

8.1 What regression R² can and cannot establish

A regression R² describes the proportion of observed outcome variance explained by the predictors in a specified dataset and model. It is not the proportion of an individual result that was ‘created’ by those predictors. It is also sensitive to the range and distribution of predictor values, outcome measurement error, omitted variables, model form and sampling. A narrow or highly structured cohort can produce a different R² from a broad clinical population even if the underlying processing system is unchanged. [2,4,10–12]

A high R² for age and weight is therefore a serious finding about dependence at population level. It may indicate strong algorithmic weighting, biological covariance, reference-architecture effects or a combination. It does not identify the mechanism without additional experimental information. Conversely, a lower R² does not prove that the residual is ultrasound-derived bone signal. Residual variation can include acquisition quality, operator technique, soft-tissue effects, disease heterogeneity and random error. [2,4,10–12]

Adjusted R², regularisation and bootstrap optimism correction address different aspects of model fit. A large decline after optimism correction indicates that the apparent model performance is unlikely to reproduce well in new data. It is not evidence that the true predictive contribution is larger than the corrected result. This distinction is particularly important in small datasets with polynomial and interaction terms, where flexible models can fit noise. [2,4,10–12]

Where researchers wish to quantify the incremental value of ultrasound-derived information, the analysis should compare nested prespecified models: demographics alone; spectral variables alone; demographics plus spectral variables; and, where appropriate, clinical variables. Performance must then be assessed in a held-out or external dataset using calibration, discrimination and prediction error, not only in-sample R². [2,4,10–12]

8.2 Sampling, segmentation and outliers

The REMS reference architecture described in the literature uses age and BMI segments. Sparse sampling within those segments is a genuine limitation when the research question concerns within-segment ranges, diagnostic outliers or whether a method identifies people whose bone status differs substantially from demographic expectation. Small cell counts make observed ranges unstable and make rare extremes easy to miss. [2,4,10–12]

That problem must be separated from the sample size required for a regression using age and weight as continuous predictors. No universal requirement of 100 observations in every algorithmic cell follows from the existence of 100 reference observations per cell. Power depends on the estimand, number of parameters, distribution, effect size, measurement error and validation strategy. The Birch observation is most persuasive as a reason to prespecify cell-based analyses and collect adequately dense strata, not as a retrospective rule declaring every smaller study invalid. [2]

Outlier validation is especially important. A technology intended for individual assessment must be tested not only near population means but in people with prior fragility fracture, treatment exposure, inflammatory disease, very low or high BMI, degenerative spinal change and discordant site results. Those analyses should use clinically meaningful comparator pathways rather than assume that disagreement with one modality automatically makes the other result wrong. [2,4,10–12]

8.3 Correlation, agreement, discrimination and monitoring

Correlation measures whether two quantities vary together; it does not show that their numerical values agree. Strong REMS–DXA correlations in the Liu meta-analysis therefore support convergent validity at population level, while wide Bland–Altman limits can simultaneously make individual substitution unsafe. [4,9] Both statements are statistically compatible and should appear together whenever comparative performance is described. [4,9]

Known-groups discrimination asks whether a method separates groups expected to differ on the construct of interest. Montatore contributes to this question. It does not provide the same evidence as a same-patient accuracy study, and it cannot show change without baseline measurements. A useful known-groups result can coexist with uncertain calibration and poor agreement with another modality. [5]

Longitudinal monitoring requires still another evidential layer. A method may discriminate groups cross-sectionally yet be insufficiently precise or stable to detect change within one person. Monitoring studies should report short- and long-term precision, least significant change, operator and equipment consistency, biological plausibility of the time course, and comparison with an external indicator where feasible. Apparent treatment-associated change smaller than established measurement error should not be interpreted as response. [6–8,23–25]

These distinctions lead to a disciplined vocabulary. REMS may correlate with DXA without being interchangeable; may distinguish groups without proving causal change; and may detect a longitudinal signal without yet having a complete monitoring standard. Precision in language protects patients and makes the research programme clearer. [4–9,23–25]

9. An independent research programme to resolve the dispute

The current disagreement can be resolved empirically. IIMHE proposes that future work should be organised around technical decomposition, independent replication, clinically relevant validation and transparent reporting rather than further rhetorical comparison of isolated R² values. [2–5,8–12]

9.1 Technical access and reproducibility

An independent technical group should receive controlled access to de-identified raw radiofrequency data, quality metrics, the osteoporosis score before integration and the demographic variables used by the software. The protocol should be agreed in advance by researchers with REMS, DXA, ultrasound physics, biostatistical and clinical expertise. Intellectual property can be protected through a secure analysis environment without preventing independent verification. [3,4,8,10,11]

The primary objective should be to reproduce the reported BMD and Fragility Score from defined inputs and quantify the incremental information contributed at each processing stage. Sensitivity analyses should vary demographic inputs while holding spectral input constant, and vary spectral inputs while holding demographics constant. The analysis should distinguish reference-selection effects from weighting within the final equation. [3,4,8,10,11]

Software version, reference-database version and any automated quality-rejection rules must be recorded. A result that changes after a software update is not necessarily wrong, but the change must be traceable and validation repeated. Comparable version control is routine for quantitative medical software and should be part of REMS governance. [3,4,8,10,11]

9.2 Required study designs

First, a large multicentre cross-sectional study should recruit a prespecified distribution across sex, age, BMI, ancestry, disease state and diagnostic category. Each site should use a common acquisition protocol, documented operator competency and central blinded quality review. DXA comparators should be quality assured and cross-calibrated where more than one system is used. Analyses should be locked before outcome access. [2,4,8,14,15,24,25]

Second, the within-stratum hypothesis raised by Birch should be tested directly. The study should report how estimates of mean, variance, range and diagnostic-category coverage stabilise as sample size increases within each age/BMI cell. Resampling can show whether approximately 100 observations are required, whether fewer are adequate for some estimands and whether the threshold differs by site or population. [2]

Third, an outlier-enriched validation study should recruit people whose bone status is likely to depart from demographic expectation: long-term glucocorticoid users, people receiving potent anabolic or antiresorptive treatment, patients with prior fragility fracture, endocrine disease and marked site discordance. The aim should be to determine whether REMS correctly identifies clinically important departures rather than merely reproducing group averages. [2,4,5,10,11]

Fourth, prospective longitudinal cohorts should include baseline and repeated REMS, repeated quality-assured DXA where appropriate, documented therapy and adherence, anthropometric stability or change, and fracture or validated intermediate outcomes. Analysis must separate short-term measurement precision from biological change and should include independent adjudication of discordant results. [4,6–8,23–25]

9.3 Clinical endpoints and population validity

The target endpoint must be defined. Agreement with DXA is relevant for compatibility with established thresholds, but DXA is not a perfect biological truth standard. Fracture prediction, treatment-associated change, artefact-resistant assessment and clinical decision impact are different endpoints. A study designed for one should not be presented as proving all the others. [4–9,13,21,26–28]

External validation should include women and men, varied ancestry, different BMI distributions and disease-specific populations. Diversity should be treated as a required validation domain rather than a nuisance. Individual subgroups must be large enough for estimates of calibration and error; simply recruiting a diverse but small total sample does not establish equitable performance. [4,5,9–11]

Harms and workflow outcomes also matter: false reassurance, unnecessary referral, inappropriate treatment escalation, delayed vertebral imaging, accessibility, patient acceptability and the consequences of discordant REMS and DXA. Clinical utility is not captured by a correlation coefficient alone. [1,4,5,8,9]

9.4 Reporting, governance and disclosure

Future publications should report operator training and experience, acquisition failures and exclusions, software version, missing data, scanner and site effects, prespecified primary analyses, all relevant conflicts of interest and the role of the manufacturer. Protocols and statistical analysis plans should be registered where feasible, and de-identified analysis datasets or executable code made available to independent reviewers. [2–5,8,11,24,25]

The same standards should apply to critical studies. Ownership of REMS or DXA services, consultancy, device-company support and institutional advocacy should be disclosed consistently. Disclosure informs interpretation; it must not be used as a shortcut for accepting or rejecting results. [2–5,10–12]

10. Specific amendments requested from MelioGuide

IIMHE asks MelioGuide to consider the following concrete amendments. These requests are intended to improve accuracy and do not require the article to adopt IIMHE’s overall interpretation. [1–5]

  • Correct the description of IIMHE’s relationship with OsteoScan UK and remove the claim that IIMHE is partnered with a REMS distributor. [1]
  • State that R² is a population-level measure of explained variance and not the percentage of an individual’s result generated by demographics. [4,10–12]
  • Present the Birch close-to-100 observation as a postulate and trend requiring validation, and report accurately that ranges exceeded 2.0 units in segments with more than 50 scans. [2]
  • Discuss the Baquir study, including its strong demographic association, wide limits of agreement, small sample, clinical heterogeneity and two uncross-calibrated DXA systems. [4]
  • Discuss Montatore as evidence of group-level discrimination after age/BMI matching, while making clear that its cross-sectional design, absent baseline and absent DXA prevent causal, accuracy or monitoring claims. [5]
  • Describe the Chan volunteer experiment as evidence that entered demographics change output, while distinguishing that finding from a quantitative decomposition of a correctly entered patient’s result. [11]
  • Frame FRAX double counting as a warning against entering REMS-BMD into FRAX, not as an error inherent in standard REMS reporting. [1,11,28]
  • Separate reference-database effects on T- and Z-scores from disagreement in measured BMD expressed in g/cm². [14–16]
  • Do not add demographic and scanner R² values as if they were independent variance components. [4]
  • Clarify that REMS B-mode is for localisation and acquisition guidance and is not diagnostic vertebral imaging or VFA; also distinguish routine AP DXA from lateral VFA. [1,8]
  • Use consistent conflict-of-interest language across REMS-associated, DXA-associated and institutional contributors without implying that a disclosed interest proves bias. [2–5,10–12]
  • Retain the conclusion that evidence is evolving and commit to updating the article as independent technical and clinical validation becomes available. [2–11]

Conclusion

The August 2026 MelioGuide update is substantially stronger than the original article. It incorporates monitoring studies, practice parameters, guideline context, meta-analytic evidence and more balanced scrutiny of TBS. IIMHE recognises those improvements and the editorial willingness behind them. [1,6–10]

Important issues nevertheless remain. The article does not yet analyse the Baquir or Montatore studies, gives insufficient attention to what Birch actually reports, and sometimes moves from evidence of demographic influence to a stronger conclusion than the available methods can support. The same concern applies in the opposite direction: within-segment variation and group discrimination cannot automatically be labelled proof of skeletal accuracy. [1,2,4,5]

IIMHE’s position is that the relative contributions of demographics, ultrasound-derived information, acquisition effects and noise have not been definitively partitioned. Bobelyak, Chan and Baquir provide consistent evidence of strong demographic association. Conversano provides a lower estimate in a much larger manufacturer-associated dataset. Birch shows substantial within-segment variability, while Montatore shows group discrimination after matching on mean age and BMI. Each contributes information; none settles the complete question. [2,4,5,10–12]

DXA has had approximately 39 years since commercial introduction to develop cross-calibration, reference-database, precision and least-significant-change standards. REMS is in an earlier phase of standardisation. That is a reason for urgency and methodological discipline, not exemption from scrutiny and not premature dismissal. [14–25]

The next step should be independently led, adequately powered, protocol-compliant research with prespecified demographic modelling, subgroup and outlier analyses, appropriate cross-calibration, access to pre-integration spectral data and validation in diverse external populations. Until then, clinical use should remain cautious, transparent and embedded in qualified interpretation. [2–5,8–15,23–25]

Declarations

Institutional position

This response is issued by the International Institute for Musculoskeletal Health Education (IIMHE). IIMHE has an educational relationship with OsteoScan UK Ltd. IIMHE states that it has no commercial relationship with Echolight or with a distributor of REMS equipment.

Relevant interests

Nick Birch is a co-owner of and holds stock in OsteoScan UK Ltd. David Tognarini is a co-owner of and holds stock in Bone Compass, Australia. These interests are declared so readers can consider them alongside the evidence and arguments presented.

Funding and data

No external funding was received for preparation of this response. No new participant-level data were generated. Statements concerning the Birch dataset refer to the peer-reviewed letter and its data-availability declaration.

Editorial responsibility

This response was prepared under the editorial oversight of the IIMHE Founders Board. Evidence and online publication status were checked to 15 August 2026. The final published text is issued by the IIMHE Founders Board.

References

  1. Martin M. REMS Bone Scan (Echolight): Is This the DEXA Alternative? MelioGuide. Updated 12 August 2026. MelioGuide article webpage
  2. Birch N, Tognarini D, McCloud P, Button P. In response to Bobelyak et al. (2025), Pocock and Chan (2025) and Chan et al. (2026). Osteoporosis International. 2026. doi:10.1007/s00198-026-08179-z.
  3. Pocock N, Bobelyak M, Vaculik J, Stepan JJ. A rebuttal letter: Author response to OSIN-D-26-00862: Letter to the Editors: REMS is not just an age and weight calculator. Osteoporosis International. 2026. doi:10.1007/s00198-026-08180-6.
  4. Baquir PJ, Au H, Green N, et al. Radiofrequency Echographic Multi Spectrometry (REMS) for the measurement of bone density at the lumbar spine and hip — a ‘real-life’ Australian validation study. Bone. 2026;212:118012. doi:10.1016/j.bone.2026.118012.
  5. Montatore M, Guglielmi R, Isaac A, et al. Radiofrequency Echographic Multi-Spectrometry for bone health assessment in prostate cancer patients receiving androgen deprivation therapy: a prospective comparative observational study. Aging Clinical and Experimental Research. 2026. doi:10.1007/s40520-026-03467-4.
  6. Semeraro A, Chiala A, Carafa A, et al. Very short-term monitoring of romosozumab longitudinal effects in a cohort of postmenopausal women by means of Radiofrequency Echographic Multi-Spectrometry (REMS) technology. Aging Clinical and Experimental Research. 2026;38:137. doi:10.1007/s40520-026-03391-7.
  7. Caffarelli C, Forcignano R, Muratore M, et al. Longitudinal monitoring of denosumab-associated bone changes by REMS in postmenopausal women with ER-positive breast cancer receiving aromatase inhibitors. Aging Clinical and Experimental Research. 2026;38:151. doi:10.1007/s40520-026-03402-7.
  8. Zambito K, Kushchayeva Y, Bush A, et al. Proposed practice parameters for the performance of radiofrequency echographic multispectrometry (REMS) evaluations. Bone & Joint Open. 2025;6(3):291–297. doi:10.1302/2633-1462.63.BJO-2024-0107.R1.
  9. Liu RY, Liu E. Correlation between radiofrequency echographic multi-spectrometry (REMS) and dual-energy X-ray absorptiometry (DXA): a systematic review and meta-analysis. Bone. 2026;206:117820. doi:10.1016/j.bone.2026.117820.
  10. Conversano F, Pisani P, Casciaro S. Methodological clarification and analysis of demographic and anthropometric determinants in the calculation of REMS bone mineral density. Calcified Tissue International. 2026;117(1):85. doi:10.1007/s00223-026-01547-1.
  11. Chan D, Chen W, Yabsley E, Pocock N. Demographic determinants of REMS-derived BMD and fragility score. Osteoporosis International. 2026. doi:10.1007/s00198-026-07960-4.
  12. Bobelyak M, Vaculik J, Stepan JJ. Bone mineral density assessment using radiofrequency echographic multispectrometry (REMS) in patients before and after total hip replacement. Osteoporosis International. 2025;36(11):2237–2244. doi:10.1007/s00198-025-07685-w.
  13. Pisani P, Conversano F, Muratore M, et al. Fragility Score: a REMS-based indicator for the prediction of incident fragility fractures at 5 years. Aging Clinical and Experimental Research. 2023;35(4):763–773. doi:10.1007/s40520-023-02358-2.
  14. Genant HK, Grampp S, Gluer CC, et al. Universal standardization for dual X-ray absorptiometry: patient and phantom cross-calibration results. Journal of Bone and Mineral Research. 1994;9(10):1503–1514. doi:10.1002/jbmr.5650091002.
  15. Fan B, Lu Y, Genant H, Fuerst T, Shepherd J. Does standardized BMD still remove differences between Hologic and GE-Lunar state-of-the-art DXA systems? Osteoporosis International. 2010. doi:10.1007/s00198-009-1062-3.
  16. Faulkner KG, Roberts LA, McClung MR. Discrepancies in normative data between Lunar and Hologic DXA systems. Osteoporosis International. 1996;6:432–436. doi:10.1007/BF01629574.
  17. Genant HK, Engelke K, Fuerst T, et al. Noninvasive assessment of bone mineral and structure: state of the art. Journal of Bone and Mineral Research. 1996;11(6):707–730. doi:10.1002/jbmr.5650110602.
  18. Carter DR, Bouxsein ML, Marcus R. New approaches for interpreting projected bone densitometry data. Journal of Bone and Mineral Research. 1992;7(2):137–145. doi:10.1002/jbmr.5650070204.
  19. Tothill P, Pye DW. Errors due to non-uniform distribution of fat in dual X-ray absorptiometry of the lumbar spine. British Journal of Radiology. 1992;65(777):807–813. doi:10.1259/0007-1285-65-777-807.
  20. Tothill P, Avenell A. Errors in dual-energy X-ray absorptiometry of the lumbar spine owing to fat distribution and soft tissue thickness during weight change. British Journal of Radiology. 1994;67(793):71–75. doi:10.1259/0007-1285-67-793-71.
  21. Marshall D, Johnell O, Wedel H. Meta-analysis of how well measures of bone mineral density predict occurrence of osteoporotic fractures. BMJ. 1996;312(7041):1254–1259. doi:10.1136/bmj.312.7041.1254.
  22. NIH Consensus Development Panel on Osteoporosis Prevention, Diagnosis, and Therapy. Osteoporosis prevention, diagnosis, and therapy. JAMA. 2001;285(6):785–795. doi:10.1001/jama.285.6.785.
  23. Lenchik L, Kiebzak GM, Blunt BA; International Society for Clinical Densitometry Position Development Panel and Scientific Advisory Committee. What is the role of serial bone mineral density measurements in patient management? Journal of Clinical Densitometry. 2002;5 Suppl:S29–S38. doi:10.1385/JCD:5:3S:S29.
  24. Baim S, Wilson CR, Lewiecki EM, et al. Precision assessment and radiation safety for dual-energy X-ray absorptiometry: position paper of the International Society for Clinical Densitometry. Journal of Clinical Densitometry. 2005;8(4):371–378. doi:10.1385/JCD:8:4:371.
  25. International Society for Clinical Densitometry. Official Adult Positions, updated 2023. ISCD official positions webpage
  26. Tomai Pitinca MD, Fortini P, Gonnelli S, Caffarelli C. Could Radiofrequency Echographic Multi-Spectrometry (REMS) overcome the limitations of BMD by DXA related to artifacts? A series of 3 cases. Journal of Ultrasound in Medicine. 2021;40:2773–2777. doi:10.1002/jum.15665.
  27. Caffarelli C, Tomai Pitinca MD, Al Refaie A, et al. Could radiofrequency echographic multispectrometry (REMS) overcome the overestimation in BMD by dual-energy X-ray absorptiometry (DXA) at the lumbar spine? BMC Musculoskeletal Disorders. 2022;23:469. doi:10.1186/s12891-022-05430-6.
  28. Kanis JA, Johnell O, Oden A, Johansson H, McCloskey E. FRAX and the assessment of fracture probability in men and women from the UK. Osteoporosis International. 2008;19:385–397. doi:10.1007/s00198-007-0543-5.

IIMHE Founders Board · 16 August 2026

You may also enjoy reading...

NEXT WEBINAR

2nd February 2026 @ 20:00 GMT

Nutrition and musculoskeletal health

Given by: Dr Kimberley Zambito, Orthopaedic Consultant 

Hosted by: Dr Nick Birch

Throughout 2026 and beyond,  IIMHE (the International Institute of Musculoskeletal Health Education), in partnership with OsteoscanUK, will host a series of engaging educational webinars, open to anyone interested in bone and musculoskeletal health.