A few of us were talking recently about what information we need to improve our care of patients with shoulder arthritis. The conversation began as a list of unanswered clinical questions and turned into a list of unanswered methodological ones. That turn is the subject of this post. The sad truth is that we cannot answer the important clinical questions using the methods we currently use.
Here is the list we started with.
1. What do we mean by "outcome"? Revision rate, final patient-reported outcome (PRO), change in PRO, attaining a threshold PRO, which PRO, satisfaction, or willingness to undergo the procedure again.
2. Which patient factors need to be controlled for? Age, sex, high or low body mass index, comorbidities, ASA score, occupation, language, resilience, expectations, home support, medical literacy, social determinants of health.
3. Which pretreatment shoulder factors need to be controlled for? Glenoid type, humeral centering, diagnosis, prior surgery, cuff status and how it was measured, bone quality, flexibility.
4. Which treatment factors need to be documented? Nonoperative management, implant type, implant position and orientation relative to what, intraoperative mobility and stability.
5. Rehabilitation: type, compliance, duration.
6. How do we control for which surgeon does the surgery? The surgeon is a large part of the method.
7. What do we consider an important difference when we compare Treatment A with Treatment B?
8. How do we handle incomplete follow-up, given that the patients still available at ten years may not have the same characteristics as the cohort we started with?
Every item on that list is a measurement problem rather than a clinical one. What follows groups them with the clinical questions they keep us from answering.
We do not know who benefits from what
Reverse shoulder arthroplasty is now used for primary osteoarthritis with an intact rotator cuff. That indication expanded from the original indication — cuff tear arthropathy — without a trial comparing it with anatomic reconstruction in comparable patients. A multicenter randomized trial of anatomic versus reverse replacement in osteoarthritis with an intact cuff is now under way in the UK, with the Shoulder Pain and Disability Index at two years as the primary outcome [1]. Until it reports, the comparison rests on observational series.
A second comparison is missing entirely. We have no non-operative or placebo-controlled arm anywhere in the shoulder arthroplasty literature. Every series in the field is a before-and-after design, and that design cannot separate the effect of the implant from natural history, regression to the mean, and the biasing effect of the consultation itself. In the CSAW trial, subacromial decompression was studied against both a surgical placebo and no treatment: neither arthroscopic arm was better than the other, with both exceeding no treatment by a margin that did not reach the minimal clinically important difference [2]. A Finnish trial reached the same result [3].
Two cautions to be aware of.
First, subacromial pain is often self-limiting and the pathoanatomy is contested; bone-on-bone glenohumeral arthritis is a structural lesion, and the improvement after arthroplasty is far larger than anything spontaneous recovery would plausibly explain.
Second, a sham arthroplasty is not ethical. However, a trial of arthroplasty versus structured non-operative care in patients with moderate radiographic disease and tolerable symptoms is feasible and has not been done. Neither has a trial of early versus delayed arthroplasty, which would at least tell us what a year of waiting costs. Instead of taking on these questions, we debate version and lateralization.
We cannot see most of our failures
Revision is the endpoint we can count, not the endpoint most patients experience. It reflects patient dissatisfaction, patient willingness to undergo another operation, surgeon willingness to offer one, and payer approval, all at once. National Joint Registry data make the gap visible: among patients with a postoperative Oxford Shoulder Score below 29, 27% of the reverse arthroplasties were in that unsatisfactory range, and less than 5% of those patients were revised, compared with 11% for anatomic total shoulder arthroplasty and 14% for hemiarthroplasty [4]. The authors read this as a relative unwillingness to revise a failed reverse. These unrevised unsatisfactory shoulders do not show up as failures if revision is the measure of failure.
Item 8 (handling incomplete follow-up) needs an additional comment. The common assumption is that the patients still available at ten years did better than those lost along the way. The direction of this survivorship bias is not established in many studies. In a single-surgeon shoulder arthroplasty series, 34.3% of patients were lost by the seventh year, and loss was not random: severe obesity, older age, and higher ASA score all raised the risk of being lost, while patients who had a complication were 43% less likely to be lost [5]. If patients with complications stay in contact more reliably, attrition may make results look worse rather than better. Loss to follow-up has mattered in arthroplasty for a long time [6], but the sign of the bias in any given cohort is rarely shown.
A related problem exists inside the survival curves themselves. Kaplan-Meier analysis treats death as censoring, which assumes the patient who died would have carried the same revision hazard as the patient who lived. In hip and knee arthroplasty, pooled cumulative incidence of revision estimated by Kaplan-Meier was 1.55 times higher (95% CI 1.43 to 1.68) than the estimate from competing-risks methods [7]. Reverse arthroplasty cohorts are older than those hip and knee cohorts, so the impact of death as censoring may be greater for us. Registries increasingly use competing-risk methods; the clinical literature mostly does not.
What we mean by "outcome" is not settled
The MCID answers a question about one patient: did this person improve enough to notice the difference? The MCID was derived by asking individual patients whether they felt better and finding the score change that matched their answer [8]. Used that way, it works. Apply it to each patient, one at a time, and record a yes or a no. Then count the yeses. That is a legitimate and useful thing to report.
What it cannot do is tell us whether two groups differ. The error is applying a threshold built for one patient to the gap between two averages.
The thresholds themselves are also less stable than their use implies. Across 39 studies of reverse arthroplasty, 87% reported MCID values, 51% reported substantial clinical benefit, 13% reported patient acceptable symptom state, and 64% took their values from a previous publication rather than calculating them; only 28% used an anchor-based method [9]. Thresholds also vary by implant type, diagnosis, and sex within a single large multicenter cohort [10]. We are comparing studies whose success criteria were derived differently, from different populations, using different anchors.
Comparing two arms within one study
This is the comparison we can actually test, because the two groups were assembled at the same time, by the same surgeons, using the same instrument and the same follow-up.
Suppose, for example, that operation A improves the average ASES score by 45 points and operation B by 39. The six-point gap is smaller than the MCID, and the usual conclusion would be that the operations are not significantly different. That conclusion does not follow. Both arms contain patients who did very well and patients who did poorly, and six points between the means tells us nothing about how many patients in each arm ended up with a shoulder they could live with.
Count instead. Set a satisfactory score before looking at the data — the National Joint Registry work used an Oxford Shoulder Score below 29 as unsatisfactory [4] — and score each patient against it. Say 78 of 100 reach it with operation A and 61 of 100 with operation B. The difference is 17 percentage points, and it is testable: a comparison of two proportions gives p = 0.009, with a 95% confidence interval of 4.5 to 29.5 points.
That interval is the important part. It says operation A is probably better but that we cannot say whether the advantage is small or large. A mean difference held against an MCID would have reported no clinically important difference.
Comparing my series with your published series
This is what we do constantly at meetings, and it cannot be tested at all.
If my series reports 78% satisfactory and yours reports 74%, there is no valid statistic to apply to that four-point gap. The two groups were never assembled to be comparable. They differ in age, diagnosis, era, and — most consequentially — in what fraction of patients came back to be counted. A p-value calculated across two papers assigns precision to a comparison that has no reasonable design behind it, yet we see those comparisons commonly: my RSA vs your RSA, my favorite implant vs your favorite implant, my TSA vs your pyrocarbon hemi.
The most we can say is descriptive: here are my outcomes, here is the definition of a good outcome I used, and here is my follow-up rate. I looked up your outcomes, definition, and follow-up rate and here they are. Readers can then judge whether their patients more closely resemble mine or yours. That is a weaker claim than a p-value, but it is the only claim the data support.
What counting does not fix
Reporting proportions does not deal with confounding. If operation A was performed on younger patients with better bone and intact cuffs, its 78% success rate is not a number an individual patient can apply to herself. Moving from "78% of that series did well" to "your chances are 78%" requires one of three things: randomization, which is what the RAPSODI-UK trial is doing [1]; adjustment or matching on the factors in items 2 and 3 of our list; or, most plainly, reporting the proportion within defined subgroups — what fraction of 55-year-old women with a B2 glenoid and an intact cuff reached a satisfactory state. The responder proportion answers the patient’s question better than a mean change does. Whether it is her number still depends on whether the patients in the series resembled her.
What this asks of us
Report the proportion of patients reaching a defined satisfactory state, alongside the mean and standard deviation. Define satisfactory in the methods, not after seeing the results. Test differences between arms of the same study, and give the confidence interval. When comparing across studies, describe the different methods and results rather than test for a significant difference. None of this requires a new study. It is arithmetic on data already in hand.
Two cautions about the instruments themselves
The first is that PROs are the right target, but they are still ceiling-limited, culturally variable, and sensitive to how the question and the preceding conversation were framed. Preferring PROs to surrogates does not rescue us if the instrument is partly measuring expectation. An absolute-state endpoint, anchored to what the patient can actually do, is closer to what we mean than any change score. The Simple Shoulder Test is one example: it asks the patient whether they can perform each of 12 functions, so the answer describes the shoulder rather than a change in it.
The second is that we have not agreed on what to measure. One partial answer exists and is under-used: an international Delphi process produced a core event set for shoulder arthroplasty, defining twelve local event groups with agreed definitions and documentation periods [11]. It covers unfavorable events rather than patient-reported outcomes, so it addresses half the problem. A consistent core functional outcome set, agreed across societies and applied consistently, would let us compare studies rather than describe them side by side.
The surgeon is an important part of the treatment
Item 6 (the surgeon is the method) is the reason most comparative questions in our field are difficult to answer as posed. A comparison of anatomic with reverse arthroplasty is not a comparison of implants. It is a comparison of implant, operator, and indication threshold together, and the operator term may be the larger one. The surgeon effect is further complicated when different surgeons contribute different numbers of patients to be analyzed.
There is a design answer that our field has largely not used. In an expertise-based randomized trial, patients are randomized to surgeons who each perform only their preferred procedure, rather than to procedures that individual surgeons perform with unequal skill and unequal conviction. The design addresses differential expertise bias and the equipoise problem at the same time, and it was proposed for surgical trials two decades ago [12]. However, this type of study would seem difficult to accomplish in practice.
Clinical problems that remain unsolved
The young patient with glenohumeral arthritis. This is the clearest gap in the field. In a systematic review of total shoulder arthroplasty in patients under 65, 17.4% had been revised at a mean of 9.4 years, 54% showed glenoid lucency, and glenoid loosening accounted for 52% of revisions; reported survivorship ranged from 60% to 80% at ten to twenty years [13]. Across 1,591 shoulders in patients under 60, no option has yet shown durable superiority [14]. Hemiarthroplasty underperforms, the anatomic glenoid has a finite life, reverse arthroplasty in this age group has little long-term data, biologic resurfacing has been largely abandoned, ream-and-run applies to a narrow group and carries an early-revision rate, and pyrocarbon has registry signals that are still immature.
Cutibacterium and the possibly-infected shoulder. We have no adequate diagnostic test, no agreed threshold for what a positive culture means, no validated prophylaxis against a dermal reservoir, and no good comparison of one-stage with two-stage revision.
Acromial and scapular spine fracture after reverse arthroplasty. A systematic review of 90 articles put the pooled rate at 2.8%, higher after primary than revision arthroplasty and higher with lateralized glenoid designs [15]. Individual series report considerably higher figures, which is itself informative about how these fractures are looked for and defined. Prediction, prevention, and treatment all remain unsettled.
Durability of reverse arthroplasty. How the construct behaves twenty years after implantation is unknown.
Technology is being adopted ahead of the evidence
Navigation, patient-specific instrumentation, robotics, and planning software all aim at a surrogate: conformity to a preoperative plan whose correctness has not itself been validated against patient reported outcomes. Two findings are worth considering side by side. A systematic review of patient-specific instrumentation found no significant difference in version error, inclination error, or positional offset compared with standard instrumentation, and reported that none of the included studies supplied patient-reported outcomes, range of motion, strength, or data on glenoid loosening [16]. A later meta-analysis of nine comparative studies found no significant difference between patient-specific and standard instrumentation in American Shoulder and Elbow Surgeons or Constant-Murley scores [17].
These data do not exclude a benefit; it’s just that published data have not yet shown one. The questions that would settle it are answerable. What degree of deviation from plan predicts a difference in patient reported outcome? What is the cost per quality-adjusted life year of each added technology? And what would the same series look like from surgeons doing modest volumes in a community hospital, the setting in which most of these operations are performed?
Three things missing from both lists
Selection into the cohort. Every list of confounders assumes the patient reached the operating room. We do not study the patients who were never offered surgery or who were offered it and declined, and those are the people who would form the comparator arm the field is missing. The denominator problem begins in clinic.
Era effects. Comparing one decade with another conflates the implant with surgical technique, anesthesia, thromboprophylaxis, outpatient pathways, physical therapy protocols, indication drift, and the version of the instrument used. Most claims that outcomes have improved are claims about a decade, attributed to a device without controlling for the other variables that are bound to have changed over the ten years.
The unit of analysis. Shoulders or patients. Bilateral cases, and whether the second shoulder’s result is independent of the first.
Where we should start
Three, in order.
The young arthritic shoulder, because it is a genuine clinical vacuum.
The visibility of clinical failure (not revision rate), because it affects every other conclusion we draw from registry and series data.
And the absence of any comparator arm, because it means we are arguing about the details of an effect whose size we have not measured.
None of these needs a new implant. They need agreement on what we are measuring, robust accounting of who is missing from the denominator, and a study design in which the surgeon is treated as part of the treatment rather than as background noise.
What arthroplasties cost
The scale of the spending is worth stating, because it sets the price of not knowing the answers.
Two published figures allow an estimate. National Inpatient Sample and National Ambulatory Surgery Sample data show that total shoulder arthroplasty in the United States rose 212% between 2012 and 2022, from 55,245 to 172,559 procedures, with incidence rising from 17.6 to 51.7 per 100,000 [18]. In a consecutive series of 1,452 primary anatomic and reverse shoulder arthroplasties at one academic institution, the mean 90-day episode-of-care cost was $25,822 for Medicare patients and $31,055 for privately insured patients [19].
Multiplying the 2022 volume by those per-case figures gives roughly $4.5 billion at the Medicare rate and $5.4 billion at the private rate. The payer mix barely matters: at 90% Medicare the figure is $4.55 billion, at 70% it is $4.73 billion, and at 60% it is $4.82 billion. Any plausible mix lands between $4.5 and $4.8 billion. That is worth stating, because payer mix is the first assumption a reader would challenge, and it turns out not to be where the uncertainty lies.
What the figures include, and what they leave out
A 90-day episode: the surgical encounter plus ninety days after. It captures the implant, the facility, personnel, physician fees, readmissions, and post-acute care within that window. It does not capture the preoperative workup — office visits, radiographs, CT scans, or three-dimensional planning. It does not capture rehabilitation or care after ninety days, revision surgery, or lost productivity. Every exclusion pushes the true figure up, so $4.5 billion is a floor rather than a measure of total spending.
The exclusion of preoperative imaging and planning deserves emphasis, because that is precisely where much of the new technology sits. The CT scan and the planning software largely fall outside the episode being measured, which means the cost of the technologies discussed above is mostly invisible in this number.
Furthermore, these numbers are out of date, and the volume trend will raise them. Three models were fitted to the same national data [18]. The logistic model, which allows for saturation, projects 228,967 procedures by 2035. The linear model projects 334,184. The log-linear model projects 905,038. Interpolating the linear model to 2026 gives roughly 222,000 procedures and $5.7 to $6.9 billion; the log-linear model gives roughly 287,000 and $7.4 to $8.9 billion. The spread between those models is far wider than any of the cost uncertainties above. We know current spending to within about ten percent, and future spending only to within a factor of two.
What the estimate is worth
The volume figure and the cost figure come from different databases, different years, and different populations, and neither study was designed to be multiplied by the other. The cost figure comes from a single academic institution between 2014 and 2020 and is not inflation-adjusted. The volume figure carries its own caution: outpatient procedures were not captured before 2016, so the reported growth rate may be somewhat overstated [18]. This is an order-of-magnitude number, not a measurement, and it is the same cross-study arithmetic this post warns against elsewhere. We offer it as a bound on the scale of the question, not as a finding.
Where the money goes
Using time-driven activity-based costing across 1,571 shoulder arthroplasties by 12 surgeons at 4 high-volume institutions, the implant accounted for 56% of episode-of-care cost for anatomic total shoulder arthroplasty and 62% for reverse, with personnel costs from check-in through the operating room accounting for a further 21% and 17% [20]. That denominator is narrower than the 90-day claims episode above, because it does not include post-acute care, so these percentages should not be applied directly to the $4.5 billion. Within the operative episode, though, the implant is the single largest line item.
Conclusion
New technologies and new implants drive these costs higher. Sorting out which of them result in improved outcomes for the patient will require more careful studies than those simply showing that patients are improved after treatment — a statement that is true for just about every method of treating shoulder arthritis, including non-operative care. Our own group examined this question directly and concluded that additional research is required to document the clinical value of these new technologies to patients with glenohumeral arthritis [21].
Codman asked the question directly over a century ago: in whose interest is it to investigate what the actual result to the patient has been? [22] Today we ask: in whose interest is it to do the hard research — common endpoints, complete follow-up, real comparison groups — that would show which treatments are better for which patients?
It seems ironic that while performing hundreds of thousands of shoulder arthroplasties and spending billions of dollars in doing so each year, we have so many unresolved problems.
—
Which is better?
Male and female Western Bluebirds, Orcas Island.
Follow on twitter/X: https://x.com/RickMatsen
Follow on facebook: https://www.facebook.com/shoulder.arthritis
Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133//
References
[1] Rodrick HL, Dias J, Watts AC, et al. Anatomic versus reverse total shoulder replacement for patients with osteoarthritis and intact rotator cuff: the RAPSODI-UK randomised controlled trial protocol. BMJ Open. 2025;15(12):e106740. doi:10.1136/bmjopen-2025-106740. PMID: 41386993.
[2] Beard DJ, Rees JL, Cook JA, et al.; CSAW Study Group. Arthroscopic subacromial decompression for subacromial shoulder pain (CSAW): a multicentre, pragmatic, parallel group, placebo-controlled, three-group, randomised surgical trial. Lancet. 2018;391(10118):329-338. doi:10.1016/S0140-6736(17)32457-1. PMID: 29169668.
[3] Paavola M, Malmivaara A, Taimela S, et al.; Finnish Subacromial Impingement Arthroscopy Controlled Trial (FIMPACT) Investigators. Subacromial decompression versus diagnostic arthroscopy for shoulder impingement: randomised, placebo surgery controlled clinical trial. BMJ. 2018;362:k2860. doi:10.1136/bmj.k2860. PMID: 30026230.
[4] O’Malley O, Davies A, Rangan A, Sabharwal S, Reilly P. Is there a difference in thresholds for revision between shoulder arthroplasty types? A National Joint Registry study. PLoS One. 2025;20(8):e0330975. doi:10.1371/journal.pone.0330975.
[5] Torrens C, MartÃnez R, Santana F. Patients lost to follow-up in shoulder arthroplasty: descriptive characteristics and reasons. Clin Orthop Surg. 2022;14(1):112-118. doi:10.4055/cios21034. PMID: 35251548.
[6] Murray DW, Britton AR, Bulstrode CJK. Loss to follow-up matters. J Bone Joint Surg Br. 1997;79-B(2):254-257. doi:10.1302/0301-620X.79B2.0790254.
[7] Lacny S, Wilson T, Clement F, Roberts DJ, Faris PD, Ghali WA, Marshall DA. Kaplan-Meier survival analysis overestimates the risk of revision arthroplasty: a meta-analysis. Clin Orthop Relat Res. 2015;473(11):3431-3442. doi:10.1007/s11999-015-4235-8. PMID: 25804881.
[8] Kamper SJ. Interpreting outcomes 3—clinical meaningfulness: linking evidence to practice. J Orthop Sports Phys Ther. 2019;49(9):677-678. doi:10.2519/jospt.2019.0705. PMID: 31475627.
[9] Yendluri A, Alexanian A, Lee AC, Megafu MN, Levine WN, Parsons BO, Kelly JD 4th, Parisien RL. The variability of MCID, SCB, PASS, and MOI thresholds for PROMs in the reverse total shoulder arthroplasty literature: a systematic review. J Shoulder Elbow Surg. 2024;33(10):2320-2332. doi:10.1016/j.jse.2024.03.051. PMID: 38754543.
[10] Simovitch RW, Elwell J, Colasanti CA, Hao KA, Friedman RJ, Flurin PH, Wright TW, Schoch BS, Roche CP, Zuckerman JD. Stratification of the minimal clinically important difference, substantial clinical benefit, and patient acceptable symptomatic state after total shoulder arthroplasty by implant type, preoperative diagnosis, and sex. J Shoulder Elbow Surg. 2024;33(9):e492-e506. doi:10.1016/j.jse.2024.01.040. PMID: 38461936.
[11] Audigé L, Schwyzer HK, Durchholz H; Shoulder Arthroplasty Core Event Set (SA CES) Consensus Panel. Core set of unfavorable events of shoulder arthroplasty: an international Delphi consensus process. J Shoulder Elbow Surg. 2019;28(11):2061-2071. doi:10.1016/j.jse.2019.07.021. PMID: 31542325.
[12] Devereaux PJ, Bhandari M, Clarke M, et al. Need for expertise based randomised controlled trials. BMJ. 2005;330(7482):88. doi:10.1136/bmj.330.7482.88. PMID: 15637373.
[13] Roberson TA, Bentley JC, Griscom JT, Kissenberth MJ, Tolan SJ, Hawkins RJ, Tokish JM. Outcomes of total shoulder arthroplasty in patients younger than 65 years: a systematic review. J Shoulder Elbow Surg. 2017;26(7):1298-1306. doi:10.1016/j.jse.2016.12.069. PMID: 28209327.
[14] Fonte H, Amorim-Barbosa T, Diniz S, Barros L, Ramos J, Claro R. Shoulder arthroplasty options for glenohumeral osteoarthritis in young and active patients (<60 years old): a systematic review. J Shoulder Elb Arthroplast. 2022;6:24715492221087014. doi:10.1177/24715492221087014. PMID: 35669623.
[15] King JJ, Dalton SS, Gulotta LV, Wright TW, Schoch BS. How common are acromial and scapular spine fractures after reverse shoulder arthroplasty? A systematic review. Bone Joint J. 2019;101-B(6):627-634. doi:10.1302/0301-620X.101B6.BJJ-2018-1187.R1. PMID: 31154841.
[16] Cabarcas BC, Cvetanovich GL, Gowd AK, Liu JN, Manderle BJ, Verma NN. Accuracy of patient-specific instrumentation in shoulder arthroplasty: a systematic review and meta-analysis. JSES Open Access. 2019;3:117-129. doi:10.1016/j.jses.2019.07.002. PMID: 31709351.
[17] Daher M, Parmar T, Boufadel P, Fares MY, Khalil W, Horneff JG, Abboud JA, Khan AZ. Patient-specific instrumentation in primary total shoulder arthroplasty: a meta-analysis of clinical outcomes. Clin Shoulder Elb. 2025;28(2):129-136. doi:10.5397/cise.2024.01095. PMID: 40340231.
[18] Heo KY, Tornberg HN, Bailey EP, Conn V, Lee JD, Gottschalk MB, Zelenski NA, Wagner ER. Evolving trends in shoulder arthroplasty: a decade of growth and future projections in comparison with hip and knee arthroplasty. J Shoulder Elbow Surg. 2026;35:2089-2098. doi:10.1016/j.jse.2026.03.019.
[19] Farronato DM, Pezzulo JD, Rondon AJ, Porrini S, McGonigal D, Getz CL, Davis DE. Effects of patient comorbidities and demographics on episode-of-care costs following total shoulder arthroplasty. J Am Acad Orthop Surg. 2023;31(9):451-457. doi:10.5435/JAAOS-D-22-00450. PMID: 36749879.
[20] Carducci MP, Mahendraraj KA, Menendez ME, Rosen I, Klein SM, Namdari S, Ramsey ML, Jawa A. Identifying surgeon and institutional drivers of cost in total shoulder arthroplasty: a multicenter study. J Shoulder Elbow Surg. 2021;30(1):113-119. doi:10.1016/j.jse.2020.04.033.
[21] Schiffman CJ, Prabhakar P, Hsu JE, Shaffer ML, Miljacic L, Matsen FA 3rd. Assessing the value to the patient of new technologies in anatomic total shoulder arthroplasty. J Bone Joint Surg Am. 2021;103(9):761-770. doi:10.2106/JBJS.20.01853. PMID: 33587515.
[22] Codman EA. The product of a hospital. Surg Gynecol Obstet. 1914;18:491-496.
The author has no financial relationships with any orthopaedic device company.










