International Institute for Musculoskeletal Health Education

IIMHE Review of: Conversano F, Pisani P, Casciaro S. Methodological Clarification and Analysis of Demographic and Anthropometric Determinants in the Calculation of REMS Bone Mineral Density. 

https://doi.org/10.1007/s00223-026-01547-1

Published 30th May 2026

This review summarises this paper for clinicians / Healthcare professionals and medical scientists who are involved in osteoporosis management and who understand bone densitometry.

Why is this paper important?

Within the last 9 months, there have been two publications questioning the ability of REMS to determine bone mineral density (BMD) and specifically suggesting there is an excessive influence of age and weight, rather than the REMS backscatter signal, on that determination (Bobelyak et al. Osteoporosis International 2025; https://doi.org/10.1007/s00198-025-07685-w. Chan et al. Osteoporosis International 2026; https://doi.org/10.1007/s00198-026-07960-4). Both papers did a small sub-study by scanning a few subjects but with very different Age/weight inputs to which the REMS system gave different BMD results. Social media users and influencers then postulated whether REMS was just a "glorified Age/Weight/BMD calculator" and does it "measure" BMD or just estimate it?

Sadly, many did not take the time to deep dive into the published papers for the technology, nor did they appreciate the 10+ yrs of R&D and the extensive validation clinical trials in over 6,000+ pts head-to-head with DXA which are all peer reviewed published. It is also worth highlighting that there are both US and EU patents that contain the details of the REMS mathematical methodology, and REMS has been government assessed and registered in the US, EU, UK, Japan, Australia, indeed over 45 countries in total.

In response to the Bobelyak paper, one editorial and several letters were published that were in favour of, or reiterated, the primary concerns. There has been a "Call for Transparency" from the manufacturer of REMS, Echolight Spa (Italy), which is reasonable, as an accurate diagnosis is critical to both patient and clinician for the accurate management of bone health. As the REMS manufacturer, Echolight has the largest dataset in the world, and it has been proposed that it is therefore their responsibility to provide this transparency. This paper is that response.

Background

It is worth reiterating several facts which are well established:

  • BMD declines as people age
  • BMD increases as body mass increases
  • Bone densitometry is difficult and must be done by trained operators.
  • DXA measures/estimates/calculates BMD by measuring the differential attenuation of two X-ray beams of different energies through bone and soft tissue, producing an areal BMD measurement (g/cm²). BMD = Bone mineral content / bone area
  • DXA is influenced by artefacts such as osteoarthritis, bone deformity or bone spurs, calcifications,
  • DXA results can also be affected by patient positioning, BMI, bone size, device settings
  • REMS (radiofrequency echographic multispectrometry) is a non-ionising ultrasound-based technology that analyses raw radiofrequency (RF) ultrasound signals acquired from the lumbar spine or proximal femur.
    • Using spectral analysis (including Fast Fourier Transform–based processing), patient-specific RF spectra are generated from the reflected ultrasound signals.
  • These patient-specific spectra are compared against reference spectral models stratified by patient characteristics including sex, age and BMI. The resulting measure (the Osteoporosis Score) is then converted through stratum-specific equations into an estimated bone mineral density (BMD).

The use of signal transformation combined with comparison against normative reference datasets is not unique to REMS. Other medical technology examples include Quantitative EEG (qEEG), pulse oximetry, ECG interpretation and elastography. The important question is not whether reference datasets are used, but whether the underlying signal contributes clinically meaningful information beyond demographic variables alone. Chan et al 2026 found that in their n=209 real-world patients, age and weight explained most of the variance for REMS-BMD (R² > 0.90). Conversano et al. set out to determine if this was a serious concern or a quirk from a low sample size. Importantly, they aimed to show how much variability in the BMD result could be attributed to the REMS backscatter echo signal and what this means in real g/cm² or diagnostic criteria.

What Conversano et al did

The researchers used 16,000 anonymous REMS scans from the manufacturer's own database, collected between 2015 and 2024. All the scans came from Caucasian women. There were 8,000 hip and 8,000 spine scans. For each body site, they split the scans into two equal halves of 4,000, a "training set" and a "test set". They used the training set half to build prediction equations and kept the test set half separate to test those equations. They first checked that the two halves were very similar in terms of age, height, weight, BMI and bone density.

They then built equations that would predict the REMS result from a patient's age and weight. They also created several more complex versions to test the limits of their statistical model when extreme circumstances were considered as Chan et al. (2026) did. To avoid being misled by a single lucky sample, they repeated the model-building 100 times, each time using a different random group of patients. This was an approach that Chan et al. (2026) did not take.

[Figure 1. Diagrammatic representation of the Analysis Conversano et al 2026 took to analyse the 16,000 REMS scans by creating a "Development" pathway to determine the equations and then apply those equations to an "Independent test" group.]

What they found

The main finding was a large gap between the Development and Independent test set results. Somewhat as expected, when the Development equations were developed and then tested on the same patient cohort, it looked very accurate (TRAIN R²). It appeared to account for up to about 93% of the differences between patients.

However, when the same equation was tested on the separate group of patients, that is the test set, it accounted for much less. Using age and weight only, it accounted for approximately 58% at the spine, 57% at the hip and 70% at the femoral neck (TEST R²). So, between 30% and 43% of the REMS result could not be explained by age and weight.

Skeletal SiteDevelopment dataset (TRAIN R²) Variance explained by age + weightUnexplained varianceIndependent Validation dataset (TEST R²) Variance explained by age + weightUnexplained varianceAbsolute drop in explainability (difference between TEST and TRAIN)
Lumbar spine (LS)77%23%58%42%−19%
Total hip (TH)83%17%57%43%−26%
Femoral neck (FN)93%7%70%30%−23%

Table 1. Variance in REMS-BMD explained by age and weight in the development and independent validation datasets using the primary anthropometric model (training set size = 400)

When equations derived from the development dataset were tested back on the same data, age and weight they appeared to explain 77–93% of REMS-BMD variability. However, when those same equations were applied to an independent validation dataset, explainability fell to 57–70%, leaving 30–43% of REMS-BMD variability unexplained. This suggests same-cohort analyses may substantially overestimate anthropometric explainability.

Further critique of the data – adding layers to Stress test the model

To further challenge the primary findings, the authors performed additional "stress testing" of anthropometric explainability beyond the simple age plus weight model. This included expanded models with

  1. additional variables and interactions,
  2. BMI-based models,
  3. Penalised regression approaches (ridge and lasso), and
  4. nonlinear spline models,

all repeatedly tested using a fixed independent validation dataset.

This is important because it addresses concerns that findings may simply reflect model simplicity, overfitting, or failure to account for more complex relationships.

ModelLumbar Spine Test (% explained)Total Hip Test (% explained)Femoral Neck Test (% explained)
Age + Weight58%57%70%
Age + BMI39%71%74%
Age + Weight + BMI64%68%75%
Expanded model66%69%75%
Ridge66%69%75%
Lasso66%69%75%
Splines68%70%75%

Table 2. Independent validation performance of progressively expanded anthropometric models for prediction of REMS-derived BMD variability. Values represent the percentage of REMS-BMD variability explained (R² ×100) in the fixed independent validation dataset at training size = 400.

Despite these increasingly permissive models, independent-test performance improved only modestly, with at least 25% of REMS-BMD variability remaining unexplained.

Overall assessment

Overall, this is a technically robust and well-powered analysis addressing an important methodological question. The central approach — deriving equations in a development cohort and repeatedly testing them in an independent validation cohort — is an appropriate method for assessing the generalizability of anthropometric explainability. The findings consistently show that very high explainability observed within development datasets (>90%) falls substantially under independent validation. On this point, the study makes a convincing case that same-cohort analyses may materially overestimate the contribution of age and body habitus to REMS-derived BMD.

However, the findings should be interpreted within the context of the study design. This was a manufacturer-led analysis using manufacturer-derived data and internal validation only. The study effectively challenges the proposition that REMS output is largely a reflection of demographic and anthropometric inputs alone. Does the remaining unexplained component represent clinically meaningful skeletal information? That residual variance could reflect patient-specific RF signal content but may also include technical or other unmeasured factors. However, since the pivotal and registration studies all showed that the REMS outputs were aligned to DXA derived BMD values clinicians should be confident that the REMS derived BMD does represent their patients' bone health status.

Again, it is very important to re-iterate the findings from the authors and this analysis.

"Using age and weight alone, approximately 30% of femoral-neck REMS-BMD variance and more than 40% of total-hip and lumbar-spine REMS-BMD variance remained unexplained". For clinicians, the practical interpretation is that these data are reassuring and provide evidence that rebuts the view that REMS functions simply as an anthropometric calculator and they show that the derived REMS output is substantially influenced by the REMS backscatter signal.

We look forward to further studies that explore the variability of results from REMS in clinical use, especially independent external validation which remains necessary to determine whether the unexplained component translates into meaningful diagnostic or prognostic value.

The bottom line is:

This is a well-conducted methodological analysis that appropriately challenges an over-stated claim, while leaving the broader clinical significance of the residual signal open for independent confirmation.

IIMHE Research Directorate
25th May 2026


Appendix

1. Understanding the statistics

This section explains the main statistical ideas behind the paper.

R-squared: the share-explained score

Almost every result in this paper is reported as a number called R-squared, usually written as R². Think of it as a score between 0 and 1, which can also be read as 0% to 100%. It answers a single question: of all the differences between patients in their REMS result, what share can this equation account for? A score of 1.00, or 100%, would mean the equation accounts for everything. A score of 0.70 means it accounts for 70%. The part it cannot account for, 30% in this case, is called the unexplained part. The paper often reports this directly, as unexplained variance, and it is simply 100% minus the R² score.

Variance: how much patients differ

Variance is a word for how spread out a set of numbers is. Some patients have dense bones, and some have less dense bones. Variance is a measure of that spread. When the paper talks about explaining variance, it means accounting for why patients differ from one another.

The key idea: testing on new data

This is the most important idea in the paper. Suppose you build an equation using one group of patients, then check how well it works on that very same group. The equation will look better than it truly is. The reason is that the equation has been shaped to fit those exact patients, including their random quirks. It is a bit like a tailor who measures one person, makes a suit to fit them exactly, then points to that perfect fit as proof the suit is well made. The honest test is whether the suit also fits a different person.

Statisticians have names for these two checks. Testing an equation on the same data used to build it is called in-sample testing. The paper also calls this the "training" result. Testing it on fresh, separate data is called out-of-sample testing, or the "test" result. The out-of-sample result is the one that shows how useful the equation really is.

Overfitting: the gap between the two scores

When an equation scores much higher in-sample than out-of-sample, the difference is called overfitting. It means the equation has partly memorised the original patients instead of learning a general rule. This paper found a large overfitting gap at every anatomical site.

For example, the femoral neck equation scored 0.93 on its own data but only 0.70 on new data. That drop is overfitting in action, and it is the central point of the whole paper.

Why the work was repeated 100 times

Splitting patients into a training group and a test group just once could give a lucky or an unlucky result by chance. To guard against this, the authors repeated the whole process 100 times. Each time, they drew a different random group of patients. They then reported the middle value, called the median, along with the lowest and highest values seen. Repeating the work like this makes the conclusion much harder to dismiss as chance. They also tried four different training-group sizes, from 100 up to 400 patients, and the results barely changed.

The different equations they tested

The paper did not rely on a single equation. It tested several, to make sure the conclusion did not depend on one particular choice.

  • The simple equation. This used just two inputs, age and weight. It is the plain version, and it matches the claim the paper most wants to test.
  • The expanded equation. This added height, squared terms and interaction terms. A squared term, such as age multiplied by itself, when plotted as a graph, lets the relationship curve instead of following a straight line. An interaction, such as age multiplied by weight, lets the effect of one input change depending on another. These extra terms give the equation more freedom to fit the data. The authors call this a maximum-explainability model, but they are honest that some of the terms have no real medical meaning. It is a stress test, not a sensible model of the body.
  • Ridge and lasso. These are versions of the method that hold the equation back on purpose. They pull its numbers toward zero so they cannot swing to extreme values. They were used to check whether the unexplained part was just a side effect of an unstable equation. It was not.
  • The spline equation. This lets the relationship follow a smooth curve of any shape. It was used to check whether the result was caused by forcing into one fixed curve shape. It was not.

Checking the two halves were a fair match

Before any of this, the authors confirmed that the training half and the test half of the patients were genuinely alike. They used a Welch's t-test, a standard check for whether two groups differ on average, along with two measures of overlap. The halves turned out to be very similar. The authors point out that this makes their study cautious rather than generous. Similar halves give the age-and-weight equations the easiest possible task, so in the messier real world the equations would likely perform worse, not better.

A note on negative scores

R-squared is normally between 0 and 1. On fresh test data, however, it can fall below zero. A negative score means the equation predicts worse than simply guessing the average value for every patient. A few of the most flexible spline equations did exactly this, which is a sign they were unstable. The authors report these awkward results openly, which is good practice.

When two inputs overlap

One extra check added BMI alongside weight. But BMI is calculated from weight and height, so the two inputs partly repeat each other. When inputs overlap heavily, the equation's numbers become unreliable. This problem is called multicollinearity, and a measure called the variance inflation factor, or VIF, shows how serious it is. Here the values were moderate. There was some overlap, but not enough to break the equations.

2. Methodological strengths

The paper has real merits, in both its design and its reporting.

  1. A large dataset. The study used 16,000 scans, with 8,000 for each body site. This is larger than the biggest prospective REMS study published so far. Bigger samples give steadier, more reliable results.
  2. A proper split between training and test data. This is the study's strongest feature. Keeping the test patients completely separate from the patients used to train the equations is the correct way to find out how an equation behaves on people it has not seen. It directly fixes the weakness the authors identify in earlier work by Bobelyak (2025) and Chan (2026).
  3. Repeating the process 100 times. By not trusting a single split of the data, the authors avoid drawing a big conclusion from one lucky or unlucky sample. Reporting the median and the full range is sound practice.
  4. Many cross-checks. The conclusion was tested with simple and complex equations, with BMI-based versions, with two held-back methods, and with a flexible curved method. When a finding survives many reasonable approaches, it is more convincing. Here, the conclusion held in every version.
  5. An honestly cautious design. Matching the two halves of the data, and using only the highest-quality scans, both make the task easier for the age-and-weight equations. This means the real-world figures are likely worse, not better. The authors are clear that their numbers are a best case.
  6. Clear and open reporting. The paper reports medians, full ranges and formal tests, and it openly describes uncomfortable results, such as the negative scores. The tables are easy to follow.
  7. A carefully framed question. The authors take care to separate two different questions: whether the REMS result is linked to age and body size, and whether that link is strong enough to make the ultrasound part unnecessary. Keeping these apart is a genuine strength, because it avoids a common confusion.

3. Methodological weaknesses

The study also has important limits. Some are openly acknowledged by the authors; others affect how far the conclusions can be trusted.

  1. A serious conflict of interest. Two of the three authors are the founder and chief executive, and the co-founder and chief technology officer, of Echolight, the company that sells REMS. Both own shares in the company. All the data come from the company's own database. The paper is, in effect, a reply that defends the product against earlier studies that questioned it. This does not make the analysis wrong, and the conflict is openly declared. But it means the wording and the emphasis should be read with care, and it makes checking by independent researchers essential. The authors themselves note that repeating the analysis with data from outside the company is still future work, which means it has not yet been done.
  2. The test was internal, not external. The test half of the data was kept separate from the training half, but both halves came from the same company database, the same software version and the same narrow group of patients. This is internal independence. It does not show how the equations would behave with a truly different group, such as patients from another country, clinic, scanner or operator. The authors acknowledge this, but it is a real limit on what the study can claim.
  3. Unexplained does not automatically mean useful. This is the most important point for a careful reader. The study shows that between a quarter and nearly a half of the REMS result is not explained by age and body size. The authors suggest this leftover part is compatible with real bone information picked up by the ultrasound. That is possible. But a leftover is only a leftover. It could be genuine bone signal, or it could be measurement noise, differences between operators, scan-to-scan variation, or quirks of the software. The authors do mention these other possibilities, but the abstract and discussion lean towards the flattering reading. The study convincingly knocks down the strong claim that REMS is just an age-and-weight calculator. It does not, by itself, prove the opposite: that the leftover part is accurate and meaningful bone measurement. Showing that would need a different study, one that compares REMS against a standard reference test such as DXA. This paper deliberately does not do that.
  4. No measure of REMS's own repeatability. A simple way to give the leftover part meaning would be to show how much a REMS result changes when the same patient is scanned twice. If REMS is highly repeatable, more of the leftover is likely to be real signal. If it is not, more of it is likely to be noise. The paper does not include this, so the reader cannot judge from it how much of the unexplained part is trustworthy. However, in previous reports the precision of REMS has been established by Di Paola et al. (Osteoporosis International 2019. https://doi.org/10.1007/s00198-018-4686-3) and it has been shown to be rather better than the published results for DXA. This adds weight to the claim in the current study that the leftover part really is ultrasound information from bone.
  5. The maximum-explainability model is not really a maximum. The expanded equation is presented as a test of the upper limit of what body measurements can explain. But the researchers only had age, height, weight and BMI to work with. They did not have body composition, meaning how much of the body is fat and how much is muscle. The DXA research the authors themselves cite found fat mass to be very important for bone density. Without it, the true ceiling for body-measurement explanation is unknown, and the word maximum claims more than was tested. The authors partly concede this point.
  6. Only Caucasian women were studied. The analysis included no men and no other ethnic groups. The findings cannot simply be assumed to apply beyond this group. This is acknowledged by the authors.
  7. Only top-quality scans were used. Everyday clinical practice includes poorer scans as well. Using only the best scans removes one source of unexplained variation, which helps the leftover part look more like real signal. But it also means the study does not reflect normal conditions. The authors list this as future work.
  8. Some persuasive framing. The paper is openly a clarification, which in practice means a rebuttal. The discussion gathers several other studies that fit its message, and at one point compares its own carefully tested results against another study's untested results in a way that favours its conclusion. The supporting evidence is reasonable, but it is assembled to make a case, and a reader should keep that in mind. Several of the cross-checks also sit in the supplementary material rather than the main paper, so the main text cannot be fully judged on its own.

IIMHE Board of Founders

You may also enjoy reading...

NEXT WEBINAR

2nd February 2026 @ 20:00 GMT

Nutrition and musculoskeletal health

Given by: Dr Kimberley Zambito, Orthopaedic Consultant 

Hosted by: Dr Nick Birch

Throughout 2026 and beyond,  IIMHE (the International Institute of Musculoskeletal Health Education), in partnership with OsteoscanUK, will host a series of engaging educational webinars, open to anyone interested in bone and musculoskeletal health.