Saturday, August 15, 2026

Recent Reports about the Reverse Total Shoulder — What Can We Learn from Them?


PART I — THE SHORT VERSION

Nine papers on the reverse total shoulder arthroplasty came across the desk in the past few weeks. Eight of them measured something other than what the patient felt about their result from the surgery. Only one asked the patient something, but only tested whether the question was reproducible.

Here is what each paper reported, and what it did not.

Study

Endpoint reported

Endpoint missing

Reverse in patients ≤65 — Obana (n = 103) [1]

Range of motion; revision

Any patient-reported outcome



Implant selection and positioning — Parsons (n = 49) [2]

Simulated impingement-free motion

Any patient-reported outcome



3-D preoperative planning — Zampetakis (635 shoulders) [3]

Version, inclination, screw placement accuracy

Evidence that 3-D planning leads to superior clinical outcomes



Scapular stress fracture risk factors — Bengart (n = 662 screened) [4]

Acromion-to-lateral-humerus distance

Any patient-reported outcome




Predicting fracture by machine learning — Schneller (n = 2,256) [5]

Fracture incidence; model discrimination

Whether the model prediction changed a decision or a result


Managing acromial and scapular spine fractures — Abukar (376 fractures) [6]

Radiographic union, by treatment

Statistical comparison of the outcome scores



Short versus standard-length stems — Hollo (n = 114) [9]

Coronal alignment; distal filling ratio

Any patient-reported outcome




Late central screw breakage — Romem (n = 3) [7]

Mode and timing of construct failure

How many similar cases in the surgeon’s practice did not fail 


Reliability of the Forgotten Joint Score — Feik (n = 60) [8]

Reliability of the FJS-12

Validity, responsiveness, reverse-specific behavior


Every endpoint in the middle column got sharper. Version error is now measured in degrees rather than described in words. Fracture risk carries a probability. Stem alignment carries a threshold. 

The column on the right is mostly empty. The largest single-center series of young patients having RSA collected no patient-reported outcomes. The most faithfully executed plan has not been shown to change what patients report about the outcome of their surgery. Fixing an acromial fracture heals the bone more often than non operative treatment, but does not make the patient significantly better. The one instrument built to hear from the patient about their outcomes has been tested in only eleven reverse shoulders.

When the accessible endpoints improve every year, yet the patient’s perception of their outcome is not collected, we can try to convince ourselves that we’re making progress, but we can do more meaningful patient outcomes research.

PART II — THE PAPERS, ONE AT A TIME

The younger patient

Obana and colleagues at Columbia [1] reported on 103 patients aged 65 or younger finding improved range of motion after RSA. While 187 patients met the age criterion, only 30 (16%) contributed the two-year motion data. Patient-reported outcome measures were not collected routinely in clinic; none are presented. Yet the abstract concludes that the surgery “can reliably improve clinical outcomes.” Degrees of elevation and a patient’s account of their shoulder are not the same. Neel and colleagues found lower satisfaction and function in younger patients despite comparable postoperative motion. Eight patients (7.8%) required revision at a mean of 25.8 months, baseplate failure was the most common indication.


Optimizing what we can simulate

Parsons and colleagues at Duke [2] ran 49 preoperative plans through 19 implant configurations each and 12 motions, using a model that includes scapulothoracic as well as glenohumeral motion. The model indicated that larger, eccentric, lateralized glenospheres produced more impingement-free motion. The authors state that the surgeries were too recent to correlate with any postoperative result, that the model captures bone-on-bone and implant-on-bone contact but not deltoid or cuff tension, that osteophytes were left in place, and that the motion patterns simulated are those of healthy subjects rather than reverse patients. 

Zampetakis and colleagues [3] conducted a review covering 11 studies and 635 reverse arthroplasties. Three-dimensional planning, with or without patient-specific instrumentation or navigation, improved the accuracy of glenoid component and peripheral screw placement with respect to a preoperative plan, with version error falling from 6.7 ± 5.4 to 1.5 ± 1.1 degrees in the largest comparative study. However, they found no evidence that patient comfort and function were improved by the use of 3D planning. Clinical outcomes were assessed in only 3 of the 11 studies. In two of them the Constant, ASES, and DASH scores were comparable to cases in which 3D planning was not used. The third study had no conventional-technique arm. So, the glenoid component placement became more accurate without evidence that the patient’s outcome improved.


The acromion

Bengart and colleagues [4] screened 662 primary reverse arthroplasties and found 14 acromial or scapular spine stress fractures, a rate of 2.1%, matched 3:1 to 42 controls. The fracture group had a postoperative acromion-to-lateral-humerus distance of 13.1 mm versus 8.5 mm (P = 0.034), and that distance increased by a median 2.3 mm from preoperative values while the control group’s decreased by 3.0 mm (P = 0.024). Receiver operating characteristic analysis identified 9.78 mm as the most discriminant cutoff, with an area under the curve of 0.690. Their practice is now to avoid lateralizing beyond 10 mm. Their data supporting this practice is thin. Two measurements are in play: (1) the absolute distance on the postoperative film, which is what the 10 mm rule is based on, and (2)  the change in that distance from before surgery to after. The authors built a separate model for each. The model using the absolute distance did not reach significance (P = 0.064). The model that did reach significance used the change, and its coefficient runs the other way: β = −0.132 per millimeter. Each additional millimeter of increase in acromion-to-lateral-humerus distance was associated with a 12% lower odds of fracture, which is opposite to the group comparison the paper draws its recommendation from. 

Schneller and colleagues [5] used 2,256 registry patients to build a logistic regression model. Acromial and scapular spine fracture incidence was 4.1%, of which 85% were Levy type 2 or 3. On the held-out test set the model achieved sensitivity 0.71, specificity 0.61, and an area under the curve of 0.71. Of the 235 shoulders in which the model predicted a fracture, only 17 actually had one and 218 did not; of the 352 shoulders predicted not to fracture, 7 had one. Implant design did not differ significantly between the fracture and no-fracture groups on straightforward comparison (P = .348). The authors did not indicate whether the use of the model changed choices made in their practice.

Abukar and colleagues [6] reviewed fourteen studies, 376 fractures in 374 patients, mean age 72.9 years, 78% women, 70% atraumatic, mean time to diagnosis 11.3 months. Union was 68% overall: 78% after fixation and 59.8% after nonoperative care, a difference that reached statistical significance (P = 0.004). The abstract and conclusion both state that functional outcomes are comparable between the two approaches. However, that claim was not tested. Four scores were reported. On three of them the operated patients did slightly better: VAS 1.7 against 2.4, ASES 66.9 against 60.2, and Subjective Shoulder Value 63.5 against 60.7. On the Constant score they did worse, 37.9 against 46.0, but only two of the fourteen studies reported a Constant score after surgery, and those two contributed only 12 patients between them. The authors state that none of these scores could be compared statistically because patient-level data were inconsistently reported.


The stem and the screw

Hollo and colleagues [9] compared coronal alignment between short and standard-length humeral stems of the same design in 114 consecutive primary reverses at two Swiss centers, 57 stems in each group, on true anteroposterior radiographs at a minimum of six months. Median varus-valgus deviation was 3.50 degrees (IQR 1.90 to 5.10) for short stems and 1.90 degrees (IQR 1.00 to 3.90) for standard-length stems (P = 0.002). Deviation beyond 5 degrees occurred in 26.3% of short stems and 12.3% of standard-length stems. Standard-length stems held alignment above an observed distal filling ratio of 48%, short stems above 59%, with an absolute risk reduction of 30.4%. The authors call this threshold analysis exploratory, note that it was derived from the observed data without receiver operating characteristic analysis, and say it has not been validated in an independent cohort. The difference between groups was a median of 1.2 degrees. The primary regression model explains 6.7% of the variance in alignment; most of what decides where a stem sit was not measured: for example, broaching trajectory, insertion force, humeral deformity, and canal flare. The clinical importance of stem alignment was not tested; no patient-reported outcome, range of motion, or pain data were collected

Romem and colleagues [7] report three patients, 61 to 81 years old, in whom a monoblock central screw broke between 5.5 months and 8 years after reverse arthroplasty performed with bone graft behind the baseplate. Two had abrupt pain and functional decline after uneventful recoveries, one of them while catching a falling door; the third developed activity-related pain three and a half years after surgery.  The authors did not report the number of patients in the practice that did not have central screw failure after a reverse total shoulder with a monoblock central screw performed with bone graft behind the baseplate. The series does establish a pattern worth recognizing: when the baseplate sits largely on graft rather than native bone, incomplete incorporation may permit cyclic toggling that loads the central screw to fatigue. Plain radiographs were the best diagnostic test; computed tomography rarely added decisive information, An abrupt decline in the outcome after a good early result deserves an x-ray.

The one paper that asks the patient to provide the outcome

Feik and colleagues [8] took the twelve-item Forgotten Joint Score into shoulder arthroplasty and tested its reproducibility in 60 patients with glenohumeral osteoarthritis. Intraclass correlation coefficients were 0.96 to 0.97.  Only 11 of the 60 patients had a reverse for osteoarthritis. The minimal detectable change at 90% confidence was 13.5 points at six months and 14.2 at one year, and 15.1 for the cohort as a whole, on a 0 to 100 scale. That is a wide band for an instrument meant to separate good results from very good ones.


Where that leaves us

That is the last few months in nine papers. The surrogate measures are getting better. The effects on the patient’s comfort and function are mostly missing.

Keeping our focus on the patient

Great horned owl

San Antonio

Follow on twitter/X: https://x.com/RickMatsen

Follow on facebook: https://www.facebook.com/shoulder.arthritis

Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133//

References

1.        Obana KK, Chen JY, Weiss DL, Luzzi AJ, Knudsen ML, Jobin CM, Levine WN. Outcomes of reverse total shoulder arthroplasty in patients ≤65 years old. J Shoulder Elbow Arthroplast. 2026;10:100041. doi:10.1016/j.jsea.2026.100041

2.        Parsons KE, Shenoy DA, Lorentz SG, Hurley ET, Navacchia A, Moverman M, Levin JM, Klifto CS. Optimization of implant selection and positioning for reverse total shoulder arthroplasty using three-dimensional computed tomography–guided simulation software with scapulothoracic motion. J Shoulder Elbow Arthroplast. 2026;10:100049. doi:10.1016/j.jsea.2026.100049

3.        Zampetakis K, Sakellaridis G, Lepetsos P, Tsiotsias A, Leonidou A. Three-dimensional preoperative planning in reverse shoulder arthroplasty: a systematic review of implant positioning accuracy and clinical outcomes. Cureus. 2026;18(7):e112080. doi:10.7759/cureus.112080

4.        Bengart JJ, Kohut KT, Haider MN, Feng L, Duquin TR. Radiographic risk factors for scapular stress fractures after reverse total shoulder arthroplasty: a case-control study. J Am Acad Orthop Surg. 2026;34(16):e2233–e2240. doi:10.5435/JAAOS-D-25-00976

5.        Schneller T, Cina A, Maggini E, Klimov A, Braun M, Pfender A, Lazaridou A, Scheibel M. Prediction of acromial and scapular spine fractures after reverse total shoulder arthroplasty using machine learning: a retrospective cohort study. J Shoulder Elbow Surg. 2026 [in press]. doi:10.1016/j.jse.2026.06.023

6.        Abukar A, Case C, Sheth U, Henry P, Nam D. Outcomes of operative and non-operative management of acromial and scapular spine fractures after reverse shoulder arthroplasty: a systematic review. J Shoulder Elbow Arthroplast. 2026 [in press]. doi:10.1016/j.jsea.2026.100081

7.        Romem R, Kalva SR, Fucich D, Perry AJ, Yao JJ, Kwon YW. Late central screw breakage of a monoblock glenoid baseplate placed with bone graft: a report of 3 cases. JBJS Case Connect. 2026;16(3):e25.00538. doi:10.2106/JBJS.CC.25.00538

8.        Feik ML, Smith AZ, Chen KK, Gehring ZA, Myers NL, Gregory JM. The reliability of the forgotten joint score for shoulder arthroplasty. J Shoulder Elbow Arthroplast. 2026;10:100059. doi:10.1016/j.jsea.2026.100059

9.        Hollo D, Soproni I, Toft F, Ateschrang A, Shirinskiy I, Blakeney WG, Bauer S. Standard-length stems require lower distal filling ratios than short stems to achieve neutral alignment in reverse total shoulder arthroplasty. JSES Int. 2026 [in press]; article 101765.

Secondary citations referenced in the text — Neel GB et al. (J Shoulder Elbow Surg. 2022;31:1803–1809), Gauci MO et al. (J Shoulder Elbow Surg. 2024;33:1771–1780), Levy JC, Anderson C, Samson A (J Bone Joint Surg Am. 2013;95:e104), and Schoch BS et al. (J Shoulder Elbow Surg. 2022;31:1647–1657) — are cited as they appear within the nine primary sources above and have not been independently retrieved.

Friday, August 14, 2026

All the failures we cannot see

What is a failure?

We often define failure of an arthroplasty as a revision. It is evident, however, that the lack of a revision does not indicate that the patient has had a successful outcome. The lack of a revision simply indicates that the surgeon was unwilling to do another operation on the patient — or that the patient did not want more surgery, even though they were unhappy about the result.

If we equate revision and failure, many clinical failures go unseen. In the UK National Joint Registry, among 21,918 patients with a recorded postoperative Oxford Shoulder Score, 26.99% of those having a reverse total shoulder arthroplasty (RSA) had an unsatisfactory result, defined as an Oxford Shoulder Score below 29. Fewer than 1 in 20 of them (4.87%) were revised. Among patients with an unsatisfactory score, the proportion revised was 10.58% after anatomic total shoulder arthroplasty (aTSA) and 13.86% after hemiarthroplasty [1]. The authors concluded that the lower revision rate after RSA may not indicate better outcomes than anatomic arthroplasty, but rather a higher threshold for revising an RSA [1]. When an RSA fails, the prospect for improving the patient’s comfort and function is often poor, so the patient keeps the implant and keeps the disability.

Furthermore, achieving the minimal clinically important difference (MCID), the substantial clinical benefit, or the patient acceptable symptom state on the ASES, SANE, SST, or VAS after shoulder arthroplasty — the common indicators of a “successful arthroplasty” — did not correlate with patient satisfaction, with willingness to undergo the operation again, or with willingness to recommend it to a friend or family member [3]. In a prospective cohort of 1,559 RSAs, the 134 patients (8.6%) who rated themselves unchanged or worse had nevertheless improved on every measured outcome [4]. Their scores went up. They were not satisfied.

On one hand, a change in a patient-reported outcome that does not reach the MCID establishes that the patient has not improved. On the other hand, a change that exceeds the MCID does not establish that the patient is satisfied. MCID thresholds give strong evidence of failure and weak evidence of success.

Thus, the usual methods do not identify the patients who find their outcomes unsatisfactory.

Inspired by the wonderful book, “All the Light We Cannot See,” we wondered about “All the Failures We Cannot See”.


We are realizing that failure needs to be defined more broadly: the patient is dissatisfied or in pain, has sustained a complication, has been revised, or is simply not better than before surgery. The 1 in 10 patients who report themselves unchanged or worse after an anatomic total shoulder arthroplasty [5], and the roughly 1 in 5 who are dissatisfied after total knee arthroplasty [6], are where we can look for opportunities to improve our method by asking, as Codman would have us do, what could we have done differently that might have prevented the failure? [2, 8, 9]. Over five years Codman tracked his patients by sending them End Result Cards at a year after surgery, asking whether they were better. For those who were not better, he tried to determine why not [10]. What deserves emphasis is his method: he asked each patient for follow-up, and he did it with a simple card. The modern version of Codman’s card is a text message or an email, sent on the anniversary of the operation. The medium has changed; the method has not. Asking is what makes the failures visible.

When we identify a patient with an unsatisfactory result, we have a unique opportunity to ask questions such as “If I had chosen a different implant, if my implant fixation had been better, might this failure have been avoided? Should I have operated on this patient at all? What is different about my care of this patient in contrast to comparable patients in my practice who did well?” As Pearl has argued, these counterfactual questions are the language of causal reasoning [11].

It is important to avoid naming the mode of failure and calling it the cause. “Glenoid component loosening” is not a cause; it is what happened. The question is what led to it — the bone quality, the deformity, the technique of bone preparation, or the seating of the component.

By analogy, consider a report that the battle was lost because the general did not arrive. The general’s absence was not the cause; it was what happened. The battle was lost because the farrier failed to place the horseshoe nails properly, the shoe came loose, the horse stumbled, the general broke his leg, and he could not lead the charge. Once recognized, the nail placement is the thing that must be fixed before the next battle. So the question for us is: what might we do to reduce the risk of glenoid component failure in the next case?

How can we see our failures?

It starts with a secure log of our own cases.

(1)  For each surgery, enter the following

Name | Medical Record No. | Date of Birth | Mobile | Email

Diagnosis | Procedure | Surgery Date

(2)  Prepare a short message — a text or an email, sent through the institution’s patient portal or another secure channel — to go to the patient at 1 year after surgery.

I am interested in knowing how you are doing after your surgery. Please reply to this message. Are you better than before? Have you had any problems? If so, please let me know about them.

(3)  Trigger the message from the surgery date. A calendar reminder and a delayed send will do it; an automated text service will do it without your having to remember. This trigger is the whole reason the log exists.

(4)  When the patient replies, add the response to the log. When the patient does not reply, consider asking the office to telephone them.

Follow-up Date | Improved? | Additional Surgery?

(5)  For each failure (not improved, additional surgery), compare the case and its treatment with similar cases that did not fail.

(6)  Ask yourself what might have been done differently — patient selection, characterization of anatomy, procedure and implant choice, technique, perioperative management, team and system factors — to avoid the failure. Ask, as Codman did, “why not?” Enter this information into the log.

(7)  Keep the log where you will read it. Before a comparable case, look back at what you wrote about the last failure of that kind.

(8)  Recognize that each patient who finds their outcome unsatisfactory is an opportunity to refine your method. Two failures of the same kind are not two instances of one thing. A dislocation after reverse total shoulder arthroplasty in one patient and a dislocation in another are different events — a different patient, a different anatomy, a different set of decisions, a different day in the operating room — and a unique opportunity to learn.

Final thoughts

The measure of success is the steady refinement of our method across a career. Each of us has the opportunity to see whether our patients are better, and to learn from those who are not. Each patient who reports that they are no better is one case we can learn from.

Some of us might prefer to avoid asking a question that carries the risk of getting a “bad news” response. Others might want to avoid calling the patient’s attention to a suboptimal outcome. However, the opposite may be true — by asking, we show the patient that we care. Furthermore, in seeing the failure we may identify a chance to remedy it.

The opportunity to identify, learn from and care for these patients is there for each of us.


Seeing by looking

Short eared owl

Skagit


Follow on twitter/X: https://x.com/RickMatsen

Follow on facebook: https://www.facebook.com/shoulder.arthritis

Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133//


References

[1] O’Malley O, Davies A, Rangan A, Sabharwal S, Reilly P. Is there a difference in thresholds for revision between shoulder arthroplasty types? A National Joint Registry study. PLoS One. 2025;20(8):e0330975. doi:10.1371/journal.pone.0330975

[2] Menendez ME, Matsen FA 3rd. Learning from surgical failures. J Bone Joint Surg Am. 2026;108(8):547-548. doi:10.2106/JBJS.25.01110

[3] Khan AZ, Vaughan A, Aman ZS, Lazarus MD, Williams GR, Namdari S. Reaching MCID, SCB, and PASS for ASES, SANE, SST, and VAS following shoulder arthroplasty does not correlate with patient satisfaction. Semin Arthroplasty JSES. 2024;34(4):819-826. doi:10.1053/j.sart.2024.03.017

[4] Parsons M, Routman HD, Roche CP, Friedman RJ. Patient-reported outcomes of reverse total shoulder arthroplasty: a comparative risk factor analysis of improved versus unimproved cases. JSES Open Access. 2019;3(3):174-178. doi:10.1016/j.jses.2019.07.004. PMID:31709358

[5] Hao KA, Hones KM, O’Keefe DS, Elwell J, Simovitch RW, Wright TW, King JJ, Schoch BS. Does the relationship between preoperative function and achievement of clinically important benchmarks of success after total shoulder arthroplasty depend on outcome assessment design? Clin Orthop Relat Res. 2025;483(3):377-395. doi:10.1097/CORR.0000000000003347. PMID:39778205

[6] Bourne RB, Chesworth BM, Davis AM, Mahomed NN, Charron KDJ. Patient satisfaction after total knee arthroplasty: who is satisfied and who is not? Clin Orthop Relat Res. 2010;468(1):57-63. doi:10.1007/s11999-009-1119-9. PMID:19844772

[8] Codman EA. The product of a hospital. Surg Gynecol Obstet. 1914;18:491-496.

[9] Reverby S. Stealing the golden eggs: Ernest Amory Codman and the science and management of medicine. Bull Hist Med. 1981;55(2):156-171. PMID:7020802

[10] Codman EA. A Study in Hospital Efficiency: As Demonstrated by the Case Report of the First Five Years of a Private Hospital. Boston, MA: Thomas Todd Co.; 1918.

[11] Pearl J, Mackenzie D. The Book of Why: The New Science of Cause and Effect. New York, NY: Basic Books; 2018.

Saturday, August 8, 2026

What Are the Biggest Problems in Treating Shoulder Arthritis Today?

A few of us were talking recently about what information we need to improve our care of patients with shoulder arthritis. The conversation began as a list of unanswered clinical questions and turned into a list of unanswered methodological ones. That turn is the subject of this post. The sad truth is that we cannot answer the important clinical questions using the methods we currently use.

Here is the list we started with.

1. What do we mean by "outcome"? Revision rate, final patient-reported outcome (PRO), change in PRO, attaining a threshold PRO, which PRO, satisfaction, or willingness to undergo the procedure again.

2. Which patient factors need to be controlled for? Age, sex, high or low body mass index, comorbidities, ASA score, occupation, language, resilience, expectations, home support, medical literacy, social determinants of health.

3. Which pretreatment shoulder factors need to be controlled for? Glenoid type, humeral centering, diagnosis, prior surgery, cuff status and how it was measured, bone quality, flexibility.

4. Which treatment factors need to be documented? Nonoperative management, implant type, implant position and orientation relative to what, intraoperative mobility and stability.

5. Rehabilitation: type, compliance, duration.

6. How do we control for which surgeon does the surgery? The surgeon is a large part of the method.

7. What do we consider an important difference when we compare Treatment A with Treatment B?

8. How do we handle incomplete follow-up, given that the patients still available at ten years may not have the same characteristics as the cohort we started with?

Every item on that list is a measurement problem rather than a clinical one. What follows groups them with the clinical questions they keep us from answering.

We do not know who benefits from what

Reverse shoulder arthroplasty is now used for primary osteoarthritis with an intact rotator cuff. That indication expanded from the original indication — cuff tear arthropathy — without a trial comparing it with anatomic reconstruction in comparable patients. A multicenter randomized trial of anatomic versus reverse replacement in osteoarthritis with an intact cuff is now under way in the UK, with the Shoulder Pain and Disability Index at two years as the primary outcome [1]. Until it reports, the comparison rests on observational series.

A second comparison is missing entirely. We have no non-operative or placebo-controlled arm anywhere in the shoulder arthroplasty literature. Every series in the field is a before-and-after design, and that design cannot separate the effect of the implant from natural history, regression to the mean, and the biasing effect of the consultation itself. In the CSAW trial, subacromial decompression was studied against both a surgical placebo and no treatment: neither arthroscopic arm was better than the other, with both exceeding no treatment by a margin that did not reach the minimal clinically important difference [2]. A Finnish trial reached the same result [3].

Two cautions to be aware of. 

First, subacromial pain is often self-limiting and the pathoanatomy is contested; bone-on-bone glenohumeral arthritis is a structural lesion, and the improvement after arthroplasty is far larger than anything spontaneous recovery would plausibly explain. 

Second, a sham arthroplasty is not ethical. However, a trial of arthroplasty versus structured non-operative care in patients with moderate radiographic disease and tolerable symptoms is feasible and has not been done. Neither has a trial of early versus delayed arthroplasty, which would at least tell us what a year of waiting costs. Instead of taking on these questions, we debate version and lateralization.

We cannot see most of our failures

Revision is the endpoint we can count, not the endpoint most patients experience. It reflects patient dissatisfaction, patient willingness to undergo another operation, surgeon willingness to offer one, and payer approval, all at once. National Joint Registry data make the gap visible: among patients with a postoperative Oxford Shoulder Score below 29, 27% of the reverse arthroplasties were in that unsatisfactory range, and less than 5% of those patients were revised, compared with 11% for anatomic total shoulder arthroplasty and 14% for hemiarthroplasty [4]. The authors read this as a relative unwillingness to revise a failed reverse. These unrevised unsatisfactory shoulders do not show up as failures if revision is the measure of failure.

Item 8 (handling incomplete follow-up) needs an additional comment. The common assumption is that the patients still available at ten years did better than those lost along the way. The direction of this survivorship bias is not established in many studies. In a single-surgeon shoulder arthroplasty series, 34.3% of patients were lost by the seventh year, and loss was not random: severe obesity, older age, and higher ASA score all raised the risk of being lost, while patients who had a complication were 43% less likely to be lost [5]. If patients with complications stay in contact more reliably, attrition may make results look worse rather than better. Loss to follow-up has mattered in arthroplasty for a long time [6], but the sign of the bias in any given cohort is rarely shown.

A related problem exists inside the survival curves themselves. Kaplan-Meier analysis treats death as censoring, which assumes the patient who died would have carried the same revision hazard as the patient who lived. In hip and knee arthroplasty, pooled cumulative incidence of revision estimated by Kaplan-Meier was 1.55 times higher (95% CI 1.43 to 1.68) than the estimate from competing-risks methods [7]. Reverse arthroplasty cohorts are older than those hip and knee cohorts, so the impact of death as censoring may be greater for us. Registries increasingly use competing-risk methods; the clinical literature mostly does not.

What we mean by "outcome" is not settled

The MCID answers a question about one patient: did this person improve enough to notice the difference? The MCID was derived by asking individual patients whether they felt better and finding the score change that matched their answer [8]. Used that way, it works. Apply it to each patient, one at a time, and record a yes or a no. Then count the yeses. That is a legitimate and useful thing to report.

What it cannot do is tell us whether two groups differ. The error is applying a threshold built for one patient to the gap between two averages.

The thresholds themselves are also less stable than their use implies. Across 39 studies of reverse arthroplasty, 87% reported MCID values, 51% reported substantial clinical benefit, 13% reported patient acceptable symptom state, and 64% took their values from a previous publication rather than calculating them; only 28% used an anchor-based method [9]. Thresholds also vary by implant type, diagnosis, and sex within a single large multicenter cohort [10]. We are comparing studies whose success criteria were derived differently, from different populations, using different anchors.

Comparing two arms within one study

This is the comparison we can actually test, because the two groups were assembled at the same time, by the same surgeons, using the same instrument and the same follow-up.

Suppose, for example, that operation A improves the average ASES score by 45 points and operation B by 39. The six-point gap is smaller than the MCID, and the usual conclusion would be that the operations are not significantly different. That conclusion does not follow. Both arms contain patients who did very well and patients who did poorly, and six points between the means tells us nothing about how many patients in each arm ended up with a shoulder they could live with.

Count instead. Set a satisfactory score before looking at the data — the National Joint Registry work used an Oxford Shoulder Score below 29 as unsatisfactory [4] — and score each patient against it. Say 78 of 100 reach it with operation A and 61 of 100 with operation B. The difference is 17 percentage points, and it is testable: a comparison of two proportions gives p = 0.009, with a 95% confidence interval of 4.5 to 29.5 points.

That interval is the important part. It says operation A is probably better but that we cannot say whether the advantage is small or large. A mean difference held against an MCID would have reported no clinically important difference.

Comparing my series with your published series

This is what we do constantly at meetings, and it cannot be tested at all.

If my series reports 78% satisfactory and yours reports 74%, there is no valid statistic to apply to that four-point gap. The two groups were never assembled to be comparable. They differ in age, diagnosis, era, and — most consequentially — in what fraction of patients came back to be counted. A p-value calculated across two papers assigns precision to a comparison that has no reasonable design behind it, yet we see those comparisons commonly: my RSA vs your RSA, my favorite implant vs your favorite implant, my TSA vs your pyrocarbon hemi.

The most we can say is descriptive: here are my outcomes, here is the definition of a good outcome I used, and here is my follow-up rate. I looked up your outcomes, definition, and follow-up rate and here they are. Readers can then judge whether their patients more closely resemble mine or yours. That is a weaker claim than a p-value, but it is the only claim the data support.

What counting does not fix

Reporting proportions does not deal with confounding. If operation A was performed on younger patients with better bone and intact cuffs, its 78% success rate is not a number an individual patient can apply to herself. Moving from "78% of that series did well" to "your chances are 78%" requires one of three things: randomization, which is what the RAPSODI-UK trial is doing [1]; adjustment or matching on the factors in items 2 and 3 of our list; or, most plainly, reporting the proportion within defined subgroups — what fraction of 55-year-old women with a B2 glenoid and an intact cuff reached a satisfactory state. The responder proportion answers the patient’s question better than a mean change does. Whether it is her number still depends on whether the patients in the series resembled her.

What this asks of us

Report the proportion of patients reaching a defined satisfactory state, alongside the mean and standard deviation. Define satisfactory in the methods, not after seeing the results. Test differences between arms of the same study, and give the confidence interval. When comparing across studies, describe the different methods and results rather than test for a significant difference. None of this requires a new study. It is arithmetic on data already in hand.

Two cautions about the instruments themselves

The first is that PROs are the right target, but they are still ceiling-limited, culturally variable, and sensitive to how the question and the preceding conversation were framed. Preferring PROs to surrogates does not rescue us if the instrument is partly measuring expectation. An absolute-state endpoint, anchored to what the patient can actually do, is closer to what we mean than any change score. The Simple Shoulder Test is one example: it asks the patient whether they can perform each of 12 functions, so the answer describes the shoulder rather than a change in it.

The second is that we have not agreed on what to measure. One partial answer exists and is under-used: an international Delphi process produced a core event set for shoulder arthroplasty, defining twelve local event groups with agreed definitions and documentation periods [11]. It covers unfavorable events rather than patient-reported outcomes, so it addresses half the problem. A consistent core functional outcome set, agreed across societies and applied consistently, would let us compare studies rather than describe them side by side.

The surgeon is an important part of the treatment

Item 6 (the surgeon is the method) is the reason most comparative questions in our field are difficult to answer as posed. A comparison of anatomic with reverse arthroplasty is not a comparison of implants. It is a comparison of implant, operator, and indication threshold together, and the operator term may be the larger one. The surgeon effect is further complicated when different surgeons contribute different numbers of patients to be analyzed.

There is a design answer that our field has largely not used. In an expertise-based randomized trial, patients are randomized to surgeons who each perform only their preferred procedure, rather than to procedures that individual surgeons perform with unequal skill and unequal conviction. The design addresses differential expertise bias and the equipoise problem at the same time, and it was proposed for surgical trials two decades ago [12]. However, this type of study would seem difficult to accomplish in practice.

Clinical problems that remain unsolved

The young patient with glenohumeral arthritis. This is the clearest gap in the field. In a systematic review of total shoulder arthroplasty in patients under 65, 17.4% had been revised at a mean of 9.4 years, 54% showed glenoid lucency, and glenoid loosening accounted for 52% of revisions; reported survivorship ranged from 60% to 80% at ten to twenty years [13]. Across 1,591 shoulders in patients under 60, no option has yet shown durable superiority [14]. Hemiarthroplasty underperforms, the anatomic glenoid has a finite life, reverse arthroplasty in this age group has little long-term data, biologic resurfacing has been largely abandoned, ream-and-run applies to a narrow group and carries an early-revision rate, and pyrocarbon has registry signals that are still immature.

Cutibacterium and the possibly-infected shoulder. We have no adequate diagnostic test, no agreed threshold for what a positive culture means, no validated prophylaxis against a dermal reservoir, and no good comparison of one-stage with two-stage revision.

Acromial and scapular spine fracture after reverse arthroplasty. A systematic review of 90 articles put the pooled rate at 2.8%, higher after primary than revision arthroplasty and higher with lateralized glenoid designs [15]. Individual series report considerably higher figures, which is itself informative about how these fractures are looked for and defined. Prediction, prevention, and treatment all remain unsettled.

Durability of reverse arthroplasty. How the construct behaves twenty years after implantation is unknown.

Technology is being adopted ahead of the evidence

Navigation, patient-specific instrumentation, robotics, and planning software all aim at a surrogate: conformity to a preoperative plan whose correctness has not itself been validated against patient reported outcomes. Two findings are worth considering side by side. A systematic review of patient-specific instrumentation found no significant difference in version error, inclination error, or positional offset compared with standard instrumentation, and reported that none of the included studies supplied patient-reported outcomes, range of motion, strength, or data on glenoid loosening [16]. A later meta-analysis of nine comparative studies found no significant difference between patient-specific and standard instrumentation in American Shoulder and Elbow Surgeons or Constant-Murley scores [17].

The issue is shown in a recent abstract, Impact of robotic-assisted surgery on enhancing stability of stemless shoulder arthroplastyTwo experiments are reported. In the first, 69 matched pairs of fresh-frozen cadaveric shoulders — 138 shoulders — underwent CT-based planning, and fourteen experienced shoulder surgeons performed a conventional guide-based resection on one side and a robotic-assisted resection on the other. Robotic resection was more accurate for inclination, retroversion, and resection height. In the second, 45 humeral specimens were scanned with a BMD phantom and 100 simulated resections per method were generated by introducing angular and positional variations said to represent conventional and robotic technique. Bone mineral density at the resection surface correlated closely with preoperative values under both, R² = 0.978 and R² = 0.995, and deviation was smaller in the robotic condition. 

It is not clear how stability of the stemless arthroplasty was assessed.

The proposed chain of reasoning has four links: the resection matches the plan, so density at the resection surface is preserved, so the stemless component is more stable, so the patient does better. Only the first link was measured. The second rests on a simulation whose two conditions differed by a deviation the investigators specified in advance, so the finding that deviation was smaller in the robotic condition is the input read back out. The third and fourth were not studied.

The two experiments were not reconciled. The surgeons cut bone in which density was not measured, and density was measured in specimens no surgeon cut. The accuracy gained in the first experiment was not connected to the density preserved in the second.

The two correlations were never compared with each other. It seems unlikely that  stemless humeral component are failing because resection-surface density correlated at 0.978 rather than 0.995.

The recommendation for future work is that robotic assistance be examined in less experienced surgeons — the one group the design excluded, since fourteen experienced shoulder surgeons performed every resection. That is where the commercial case for the technology actually rests, and it remains untested.

These examples do not exclude a benefit; it's just that published data have not yet shown one. The questions that would settle it are answerable. What degree of deviation from plan predicts a difference in patient reported outcome? What is the cost per quality-adjusted life year of each added technology? And what would the same series look like from surgeons doing modest volumes in a community hospital, the setting in which most of these operations are performed?


Three things missing from both lists

Selection into the cohort. Every list of confounders assumes the patient reached the operating room. We do not study the patients who were never offered surgery or who were offered it and declined, and those are the people who would form the comparator arm the field is missing. The denominator problem begins in clinic.

Era effects. Comparing one decade with another conflates the implant with surgical technique, anesthesia, thromboprophylaxis, outpatient pathways, physical therapy protocols, indication drift, and the version of the instrument used. Most claims that outcomes have improved are claims about a decade, attributed to a device without controlling for the other variables that are bound to have changed over the ten years.

The unit of analysis. Shoulders or patients. Bilateral cases, and whether the second shoulder’s result is independent of the first.

Where we should start

Three, in order.

The young arthritic shoulder, because it is a genuine clinical vacuum.

The visibility of clinical failure (not revision rate), because it affects every other conclusion we draw from registry and series data.

And the absence of any comparator arm, because it means we are arguing about the details of an effect whose size we have not measured.

None of these needs a new implant. They need agreement on what we are measuring, robust accounting of who is missing from the denominator, and a study design in which the surgeon is treated as part of the treatment rather than as background noise.

What arthroplasties cost

The scale of the spending is worth stating, because it sets the price of not knowing the answers.

Two published figures allow an estimate. National Inpatient Sample and National Ambulatory Surgery Sample data show that total shoulder arthroplasty in the United States rose 212% between 2012 and 2022, from 55,245 to 172,559 procedures, with incidence rising from 17.6 to 51.7 per 100,000 [18]. In a consecutive series of 1,452 primary anatomic and reverse shoulder arthroplasties at one academic institution, the mean 90-day episode-of-care cost was $25,822 for Medicare patients and $31,055 for privately insured patients [19].

Multiplying the 2022 volume by those per-case figures gives roughly $4.5 billion at the Medicare rate and $5.4 billion at the private rate. The payer mix barely matters: at 90% Medicare the figure is $4.55 billion, at 70% it is $4.73 billion, and at 60% it is $4.82 billion. Any plausible mix lands between $4.5 and $4.8 billion. That is worth stating, because payer mix is the first assumption a reader would challenge, and it turns out not to be where the uncertainty lies.

What the figures include, and what they leave out

A 90-day episode: the surgical encounter plus ninety days after. It captures the implant, the facility, personnel, physician fees, readmissions, and post-acute care within that window. It does not capture the preoperative workup — office visits, radiographs, CT scans, or three-dimensional planning. It does not capture rehabilitation or care after ninety days, revision surgery, or lost productivity. Every exclusion pushes the true figure up, so $4.5 billion is a floor rather than a measure of total spending.

The exclusion of preoperative imaging and planning deserves emphasis, because that is precisely where much of the new technology sits. The CT scan and the planning software largely fall outside the episode being measured, which means the cost of the technologies discussed above is mostly invisible in this number.

Furthermore, these numbers are out of date, and the volume trend will raise them. Three models were fitted to the same national data [18]. The logistic model, which allows for saturation, projects 228,967 procedures by 2035. The linear model projects 334,184. The log-linear model projects 905,038. Interpolating the linear model to 2026 gives roughly 222,000 procedures and $5.7 to $6.9 billion; the log-linear model gives roughly 287,000 and $7.4 to $8.9 billion. The spread between those models is far wider than any of the cost uncertainties above. We know current spending to within about ten percent, and future spending only to within a factor of two.

What the estimate is worth

The volume figure and the cost figure come from different databases, different years, and different populations, and neither study was designed to be multiplied by the other. The cost figure comes from a single academic institution between 2014 and 2020 and is not inflation-adjusted. The volume figure carries its own caution: outpatient procedures were not captured before 2016, so the reported growth rate may be somewhat overstated [18]. This is an order-of-magnitude number, not a measurement, and it is the same cross-study arithmetic this post warns against elsewhere. We offer it as a bound on the scale of the question, not as a finding.

Where the money goes

Using time-driven activity-based costing across 1,571 shoulder arthroplasties by 12 surgeons at 4 high-volume institutions, the implant accounted for 56% of episode-of-care cost for anatomic total shoulder arthroplasty and 62% for reverse, with personnel costs from check-in through the operating room accounting for a further 21% and 17% [20]. That denominator is narrower than the 90-day claims episode above, because it does not include post-acute care, so these percentages should not be applied directly to the $4.5 billion. Within the operative episode, though, the implant is the single largest line item.

Conclusion

New technologies and new implants drive these costs higher. Sorting out which of them result in improved outcomes for the patient will require more careful studies than those simply showing that patients are improved after treatment — a statement that is true for just about every method of treating shoulder arthritis, including non-operative care. Our own group examined this question directly and concluded that additional research is required to document the clinical value of these new technologies to patients with glenohumeral arthritis [21].

Codman asked the question directly over a century ago: in whose interest is it to investigate what the actual result to the patient has been? [22] Today we ask: in whose interest is it to do the hard research — common endpoints, complete follow-up, real comparison groups — that would show which treatments are better for which patients?

It seems ironic that while performing hundreds of thousands of shoulder arthroplasties and spending billions of dollars in doing so each year, we have so many unresolved problems.


We should conclude with a big "hat's off" to the remarkable, unprecedented randomized controlled trial launched by our colleagues in the UK. [1] Here is the abstract of their "Anatomic versus reverse total shoulder replacement for patients with osteoarthritis and intact rotator cuff: the RAPSODI-UK randomised controlled trial protocol."

Introduction: Shoulder osteoarthritis most commonly affects older adults, causing pain, reduced function and quality of life. Total shoulder replacements (TSRs) are indicated once other non-surgical options no longer provide adequate pain relief. Two main types of TSRs are widely used: anatomic TSR (aTSR) and reverse TSR (rTSR). It is not clear whether one TSR type provides better short- or long-term outcomes for patients, and which, if either, is more cost-effective for the National Health Service (NHS).

Methods and analysis: RAPSODI-UK is a multi-centre, pragmatic, two-parallel arm, superiority randomised controlled trial comparing the clinical- and cost-effectiveness of aTSR versus rTSR for adults aged 60+ with a primary diagnosis of osteoarthritis, an intact rotator cuff and bone stock suitable for TSR. Participants in both arms of the trial will receive usual post-operative rehabilitation. We aim to recruit 430 participants from approximately 28 NHS sites across the UK. The primary outcome is the Shoulder Pain and Disability Index (SPADI) at 2 years post-randomisation. Outcomes will be collected at 3, 6, 12, 18 and 24 months after randomisation. Secondary outcomes include the pain and function subscales of the SPADI, the Oxford Shoulder Score, health-related quality of life (EQ-5D-5L), complications, range of movement and strength, revisions and mortality. The between-group difference in the primary outcome will be derived from a constrained longitudinal data analysis model. We will also undertake a full health economic evaluation and conduct qualitative interviews to explore perceptions of acceptability of the two types of TSR and experiences of recovery with a sample of participants.

This effort recalls Star Trek: "Space: the final frontier. These are the voyages of the starship Enterprise. Its five-year mission: to explore strange new worlds; to seek out new life and new civilizations; to boldly go where no man has gone before!"


We wish them a safe voyage and eagerly await their findings when they land in two years.

Which is better?


Male and female Western Bluebirds, Orcas Island.

Follow on twitter/X: https://x.com/RickMatsen


Follow on facebook: https://www.facebook.com/shoulder.arthritis

Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133//

References

[1] Rodrick HL, Dias J, Watts AC, et al. Anatomic versus reverse total shoulder replacement for patients with osteoarthritis and intact rotator cuff: the RAPSODI-UK randomised controlled trial protocol. BMJ Open. 2025;15(12):e106740. doi:10.1136/bmjopen-2025-106740. PMID: 41386993.

[2] Beard DJ, Rees JL, Cook JA, et al.; CSAW Study Group. Arthroscopic subacromial decompression for subacromial shoulder pain (CSAW): a multicentre, pragmatic, parallel group, placebo-controlled, three-group, randomised surgical trial. Lancet. 2018;391(10118):329-338. doi:10.1016/S0140-6736(17)32457-1. PMID: 29169668.

[3] Paavola M, Malmivaara A, Taimela S, et al.; Finnish Subacromial Impingement Arthroscopy Controlled Trial (FIMPACT) Investigators. Subacromial decompression versus diagnostic arthroscopy for shoulder impingement: randomised, placebo surgery controlled clinical trial. BMJ. 2018;362:k2860. doi:10.1136/bmj.k2860. PMID: 30026230.

[4] O’Malley O, Davies A, Rangan A, Sabharwal S, Reilly P. Is there a difference in thresholds for revision between shoulder arthroplasty types? A National Joint Registry study. PLoS One. 2025;20(8):e0330975. doi:10.1371/journal.pone.0330975.

[5] Torrens C, Martínez R, Santana F. Patients lost to follow-up in shoulder arthroplasty: descriptive characteristics and reasons. Clin Orthop Surg. 2022;14(1):112-118. doi:10.4055/cios21034. PMID: 35251548.

[6] Murray DW, Britton AR, Bulstrode CJK. Loss to follow-up matters. J Bone Joint Surg Br. 1997;79-B(2):254-257. doi:10.1302/0301-620X.79B2.0790254.

[7] Lacny S, Wilson T, Clement F, Roberts DJ, Faris PD, Ghali WA, Marshall DA. Kaplan-Meier survival analysis overestimates the risk of revision arthroplasty: a meta-analysis. Clin Orthop Relat Res. 2015;473(11):3431-3442. doi:10.1007/s11999-015-4235-8. PMID: 25804881.

[8] Kamper SJ. Interpreting outcomes 3—clinical meaningfulness: linking evidence to practice. J Orthop Sports Phys Ther. 2019;49(9):677-678. doi:10.2519/jospt.2019.0705. PMID: 31475627.

[9] Yendluri A, Alexanian A, Lee AC, Megafu MN, Levine WN, Parsons BO, Kelly JD 4th, Parisien RL. The variability of MCID, SCB, PASS, and MOI thresholds for PROMs in the reverse total shoulder arthroplasty literature: a systematic review. J Shoulder Elbow Surg. 2024;33(10):2320-2332. doi:10.1016/j.jse.2024.03.051. PMID: 38754543.

[10] Simovitch RW, Elwell J, Colasanti CA, Hao KA, Friedman RJ, Flurin PH, Wright TW, Schoch BS, Roche CP, Zuckerman JD. Stratification of the minimal clinically important difference, substantial clinical benefit, and patient acceptable symptomatic state after total shoulder arthroplasty by implant type, preoperative diagnosis, and sex. J Shoulder Elbow Surg. 2024;33(9):e492-e506. doi:10.1016/j.jse.2024.01.040. PMID: 38461936.

[11] Audigé L, Schwyzer HK, Durchholz H; Shoulder Arthroplasty Core Event Set (SA CES) Consensus Panel. Core set of unfavorable events of shoulder arthroplasty: an international Delphi consensus process. J Shoulder Elbow Surg. 2019;28(11):2061-2071. doi:10.1016/j.jse.2019.07.021. PMID: 31542325.

[12] Devereaux PJ, Bhandari M, Clarke M, et al. Need for expertise based randomised controlled trials. BMJ. 2005;330(7482):88. doi:10.1136/bmj.330.7482.88. PMID: 15637373.

[13] Roberson TA, Bentley JC, Griscom JT, Kissenberth MJ, Tolan SJ, Hawkins RJ, Tokish JM. Outcomes of total shoulder arthroplasty in patients younger than 65 years: a systematic review. J Shoulder Elbow Surg. 2017;26(7):1298-1306. doi:10.1016/j.jse.2016.12.069. PMID: 28209327.

[14] Fonte H, Amorim-Barbosa T, Diniz S, Barros L, Ramos J, Claro R. Shoulder arthroplasty options for glenohumeral osteoarthritis in young and active patients (<60 years old): a systematic review. J Shoulder Elb Arthroplast. 2022;6:24715492221087014. doi:10.1177/24715492221087014. PMID: 35669623.

[15] King JJ, Dalton SS, Gulotta LV, Wright TW, Schoch BS. How common are acromial and scapular spine fractures after reverse shoulder arthroplasty? A systematic review. Bone Joint J. 2019;101-B(6):627-634. doi:10.1302/0301-620X.101B6.BJJ-2018-1187.R1. PMID: 31154841.

[16] Cabarcas BC, Cvetanovich GL, Gowd AK, Liu JN, Manderle BJ, Verma NN. Accuracy of patient-specific instrumentation in shoulder arthroplasty: a systematic review and meta-analysis. JSES Open Access. 2019;3:117-129. doi:10.1016/j.jses.2019.07.002. PMID: 31709351.

[17] Daher M, Parmar T, Boufadel P, Fares MY, Khalil W, Horneff JG, Abboud JA, Khan AZ. Patient-specific instrumentation in primary total shoulder arthroplasty: a meta-analysis of clinical outcomes. Clin Shoulder Elb. 2025;28(2):129-136. doi:10.5397/cise.2024.01095. PMID: 40340231.

[18] Heo KY, Tornberg HN, Bailey EP, Conn V, Lee JD, Gottschalk MB, Zelenski NA, Wagner ER. Evolving trends in shoulder arthroplasty: a decade of growth and future projections in comparison with hip and knee arthroplasty. J Shoulder Elbow Surg. 2026;35:2089-2098. doi:10.1016/j.jse.2026.03.019.

[19] Farronato DM, Pezzulo JD, Rondon AJ, Porrini S, McGonigal D, Getz CL, Davis DE. Effects of patient comorbidities and demographics on episode-of-care costs following total shoulder arthroplasty. J Am Acad Orthop Surg. 2023;31(9):451-457. doi:10.5435/JAAOS-D-22-00450. PMID: 36749879.

[20] Carducci MP, Mahendraraj KA, Menendez ME, Rosen I, Klein SM, Namdari S, Ramsey ML, Jawa A. Identifying surgeon and institutional drivers of cost in total shoulder arthroplasty: a multicenter study. J Shoulder Elbow Surg. 2021;30(1):113-119. doi:10.1016/j.jse.2020.04.033.

[21] Schiffman CJ, Prabhakar P, Hsu JE, Shaffer ML, Miljacic L, Matsen FA 3rd. Assessing the value to the patient of new technologies in anatomic total shoulder arthroplasty. J Bone Joint Surg Am. 2021;103(9):761-770. doi:10.2106/JBJS.20.01853. PMID: 33587515.

[22] Codman EA. The product of a hospital. Surg Gynecol Obstet. 1914;18:491-496.

The author has no financial relationships with any orthopaedic device company.