Part One — The question
1. What patients actually ask
Patient-reported outcomes are often reported as a single number — ASES, SANE, VAS, SAS. In fifty plus years of practice I have never had a patient with shoulder arthritis ask whether a proposed surgery is likely to give them an American Shoulder and Elbow Surgeons (ASES) score of 80, or whether their pre to postoperative change will exceed the minimal clinically important difference (MCID). I have had a great many ask whether the surgery can be expected to help them recover the ability to sleep (or to fasten a bra, tuck in a shirt, lift a grandchild, throw a ball, or do their work). These are the kind of specific outcomes that individuals with shoulder arthritis care about.
The gap between what surgeons report and what patients ask for was identified a while back. Writing in 1969 about the timing of surgery on the rheumatoid hand, Barron set out ten guidelines for judging whether and when to operate. The tenth reads that the assessment of disability should not be measured in degrees of movement but in the ability or the inability of the patient to do what they want to do [1].
We have since tested that claim in our own patients. In 74 men and 30 women with glenohumeral osteoarthritis seen before arthroplasty, active abduction was measured with an observer-independent motion capture system and compared with the Simple Shoulder Test. Abduction accounted for only 29% of the variation in the total score for women and 25% for men [2]. Function by function, the difference in abduction between shoulders that did and did not allow the patient to perform an item was significant for only 4 of the 12 items in women and 5 of 12 in men [2].
The contralateral shoulders of the same patients tell the other half of the story. There, abduction accounted for 54% and 46% of the variation, and the difference in abduction between shoulders that did and did not allow each function was significant for 10 of 12 items in women and all 12 in men [2].
Motion predicts function in a shoulder without arthritis and largely stops predicting it in a shoulder with arthritis. Other shoulder factors are at work.
2. What patients have lost by the time they reach us
Among 931 consecutive patients electing shoulder arthroplasty in our practice, the average preoperative Simple Shoulder Test score was 3.6 of 12, and the commonest single score was 1 [3]. More important than the composite score is that the three functions most often lacking were sleeping comfortably, washing the back of the opposite shoulder, and throwing overhand: only 9.6%, 10.2% and 5.9% of these patients could do them before surgery [3].
Nine in ten patients arriving for shoulder arthroplasty could not sleep comfortably. That is the tipping point. That is what brought them into the office.
The same picture appears in 544 concurrent cases. Of 281 patients having a total shoulder, only 22 could sleep comfortably — 7.8% [4].
The tipping point at which a patient decides to proceed with surgery varies with the individual and the procedure they are considering. For the 931 patients, the median preoperative score was 5 for those choosing a ream-and-run, 3 for total shoulder, and 1 for reverse [3]. Men elected surgery at a median of 4, women at 2 [3]. Preoperative scores (tipping points) were higher for younger patients, for those in healthier ASA classes, for the married, for those with commercial insurance, and for those whose shoulder problem was not work-related [3].
So patients arrive having lost different functions, at different points, for different reasons.
Part Two — We started out looking at individual functions and then migrated to lumping the different elements into a single “score”
3. 1987 — five functions, reported one at a time
Our first published series of total shoulder arthroplasty reported 50 Neer-II replacements in 44 patients at an average of 3.5 years [5]. It used the American Shoulder and Elbow Surgeons evaluation form as it then stood: it reported five activities of daily living individually, before and after surgery, for each shoulder [5].
Function | Before surgery | After surgery |
Sleep on the affected side | 4 of 50 (8%) | 43 of 50 (86%) |
Wash the opposite axilla | 15 of 50 (30%) | 42 of 50 (84%) |
Perineal care | 8 of 50 (16%) | 40 of 50 (80%) |
Comb the hair | 6 of 50 (12%) | 38 of 50 (76%) |
Use the hand at shoulder level | 4 of 50 (8%) | 34 of 50 (68%) |
Table 1. Ability to perform five activities of daily living, 50 total shoulder arthroplasties, 1987 [5]. Ability here means normal performance or mild compromise on a five-point scale, not the yes-or-no format of the later Simple Shoulder Test. Seven shoulders (14%) could perform all five before surgery; thirty-nine (78%) could afterward [5].
Two features of that paper are worth noticing. The patient’s own verdict was reported separately: 34 shoulders (68%) much better, 13 (26%) better, 3 unchanged [5]. And the evaluation form used in that paper listed fifteen functions — back pocket, perineal care, wash opposite axilla, eat with a utensil, comb hair, hand at shoulder level, carry 10 to 15 pounds, dress, sleep on the affected side, pulling, hand overhead, throwing, lifting, usual work, usual sport [5].
There was no composite score reported in the paper.
4. 1993 and 1994 — the Simple Shoulder Test, an inventory of functions
The Simple Shoulder Test appeared in print in 1993, in Chapter 32 of The Shoulder: A Balance of Mobility and Stability [18], and again the following year in Chapter 1 of Practical Evaluation and Management of the Shoulder [28]. The chapter sets out the reasoning before it sets out the questions. A diagnosis of instability, cuff disease, arthritis or frozen shoulder does not by itself indicate a need for treatment; the need arises from the effect of the condition on the patient's function, and the success of treatment is best measured by its ability to restore function [18]. The most important and practical assessment of a shoulder's function is the patient's view of it [18].
The twelve questions were not derived from a theory of shoulder mechanics. They are a minimal data set of yes-or-no questions drawn from Neer's evaluation, the ASES evaluation, and the complaints heard most often from patients presenting to the University of Washington Shoulder Team [18]. The list came out of the clinic.
The chapter is explicit about what to do with the answers. No score is derived, and results are not classified into fair, good, excellent and limited-goals categories. Instead the specific functional deficits for a given disorder, and the observed improvement in those functions after a specific treatment in the surgeon's practice, are explained to the patient in simple terms, so that consent is better informed [18]. It then gives the worked example: the patient may learn that 90% of patients having procedure P, for diagnosis D, performed by Dr. C, have a high likelihood of regaining their ability to sleep on the affected side [18]. That is procedure-specific, diagnosis-specific and surgeon-specific reporting of a single named function, published 33 years ago.
Before the test was put to clinical use, we checked the ceiling. Eighty subjects aged 60 to 70 (not patients, but members of the congregation of the church attended by Doug Harryman), with shoulders normal by history, physical examination and ultrasound, answered the twelve questions. Eighty of 80 could sleep comfortably. Eighty of 80 could also tuck in a shirt, place a hand behind the head, put a coin on a shelf at shoulder level, lift a pint to shoulder level, carry twenty pounds, toss a softball underhand and wash the opposite shoulder. Seventy-nine of 80 could lift a gallon to head level, and 77 of 80 could throw a softball overhand [28]. Twelve of twelve is what an unaffected shoulder of that age can do.
The 1994 chapter also reports how stable the individual answers are, which matters if functions are going to be reported one at a time. Seventy patients with abnormal tests were retested 5 to 30 days later, an average of 14 days. Function by function, the number giving the same answer ran from 62 of 70 for lifting a pint to shoulder level to 67 of 70 for comfort at the side, sleeping comfortably and washing the opposite shoulder [28]. Overall, 63 percent of patients answered every question identically on retest and 90 percent differed on no more than one [28]. The residual variation is not treated as a deficiency of the test; it is attributed to an actual day-to-day variation in some patients’ view of their shoulder function [28].
The way that chapter displays a surgical result is the thing worth looking at, because it is what the field stopped doing. It reports total shoulder arthroplasty for degenerative joint disease as separate functions rather than as a score, each shown before surgery and at three months, six months, and one to two years afterward. The reader sees, for each function on its own, the proportion of patients able to perform it at each stage. The caption of Figure 1-8 states what such a display gives the operating surgeon and their prospective patients three key bits of information: the typical preoperative state of patients having shoulder arthroplasty for that diagnosis, the likelihood of regaining a given function after surgery, and the recovery time for each function [28]. No composite score appears anywhere in the chapter's presentation of results.
Figure 1. Functional outcome of total shoulder arthroplasty for degenerative joint disease, reported one function at a time. Each row is a single Simple Shoulder Test function; each bar is the percentage of patients answering yes at that interval. Preoperative n = 29, three months n = 8, six months n = 15, one to two years n = 9 — separate cross-sections, not the same patients followed forward. Question 12, working full-time at a regular job, was not included. Redrawn from Figure 1-8 of [28]; the same data appear as Fig. 7, p. 511 of [18], where the last interval is labeled one to 1.5 years and the group sizes are given. Values read from the printed bars and rounded to the nearest whole patient.
This is where the SST puts the burden — on the individual surgeon. Outcomes for different surgeons using apparently identical procedures are often not the same, and the surgeon is the critical determinant of the procedure and its outcome. The surgeon is the method. Each surgeon should therefore document the functional outcomes of his or her own procedures rather than assume that their results will match anyone else's.
The Simple Shoulder Test is not a score. It is a list of the functional assets of a shoulder. Summing those assets into a single number obscures its value and implies that each of them has the same worth to the patient. That is like reporting a person's belongings as combining the car, the wristwatch, the bowling ball, the coffee maker and the Bengal cat. When we use the total number of functions a patient can perform as a score, we are using it as convenient shorthand for analysis, recognizing that it loses the granularity of reporting the ability to perform each function.
One more thing, from the form itself. As printed in the 1994 chapter, the Simple Shoulder Test does not stop at the twelve functions. Below them the form asks whether there are other important things the patient cannot do as a result of the shoulder problem, and then asks about previous doctors, tests, nonmedical treatments, injections and surgeries, anything else about the shoulder we should know, and any family history of shoulder problems [28]. The blank space in which a patient names their own lost function is on the form; this is the space in which they tell us the rest.
5. 1994 — the American Shoulder and Elbow Surgeons score, a number
In the same year, the ASES Research Committee published the standardized assessment form [6]. Ten activities of daily living, each scored 0 to 3, summed to a cumulative index out of 30. The shoulder score was then derived: the visual analog pain score contributing 50% and the cumulative activities index contributing 50%, for a total out of 100 [6]. Item 2 on that list is sleeping on the painful or affected side [6].
Work the arithmetic. Each activity contributes at most 3 points to an index of 30, and that index is scaled to half of a 100-point score. Each of the ten activities is therefore worth a maximum of five points out of one hundred. Sleep is 5% of the ASES score. Ninety percent of our patients arrive having lost it.
Figure 2. The composition of the 100-point ASES shoulder score, drawn from the scoring rules of the 1994 standardized assessment form [6]. Sleep on the affected side is one of ten activities of daily living sharing fifty points.
The form was built to facilitate communication between investigators, permit multicenter trials, and allow outcome data to be conveyed to administrators and the public [6]. Enlightening the patient was not on that list.
A weighting chosen so that no single activity dominates a summary statistic guarantees that the specific functional limitations the patient came in about cannot be singled out. Two documents, twelve months apart, made opposite choices about the same information.
6. 2001 and 2002 — functions gained and lost after arthroplasty
Fourteen years after the first series, we reported 128 consecutive total shoulder arthroplasties performed for primary glenohumeral osteoarthritis by one surgeon, with Simple Shoulder Test data on 102 (80%) at thirty to sixty months [7]. The paper opens by stating that numerical scores and improvements on various scales carry little meaning for patients [7].
It then does something that, as far as I can find, has not been repeated for this operation. For each of the twelve functions it reports the transition in both directions — of the patients who could not do a function preoperatively, how many regained it; and of the patients who could do it before surgery, how many lost it.
Function | Could not do it before | Regained it | Could do it before / lost it |
Arm comfortable at side | 31 | 100% | 71 / 6% |
Place coin on shelf | 43 | 91% | 59 / 5% |
Lift 1 lb to shoulder | 54 | 91% | 48 / 8% |
Hand behind head | 74 | 90% | 28 / 7% |
Sleep comfortably | 93 | 87% | 9 / 0% |
Tuck in shirt | 77 | 84% | 25 / 12% |
Wash contralateral back | 94 | 77% | 8 / 12% |
Lift 8 lb to shoulder | 86 | 69% | 16 / 19% |
Work full time | 63 | 64% | 39 / 5% |
Carry 20 lb at side | 34 | 59% | 68 / 4% |
Toss softball overhand | 98 | 59% | 4 / 25% (1 of 4) |
Toss softball underhand | 50 | 2% (NS) | 52 / 0% |
Table 2. Within-person transitions for the twelve Simple Shoulder Test functions, 102 total shoulder arthroplasties at 30 to 60 months [7]. Counts are of the 102 shoulders. Elven of the twelve reached P < 0.01; underhand throwing did not. The 25% loss figure for overhand throwing is one of four shoulders.
Read the sleep row. Ninety-three of 102 could not sleep comfortably before surgery, 87% of them regained it, and not one of the nine who could sleep before surgery could not sleep afterwards [7]. Sleep is the most commonly lost function in patients with shoulder arthritis, and is among the most reliably regained.
The two summary figures answer the patient’s questions directly. Of the 797 functions absent across the 102 shoulders before surgery, 582 were regained — 73% [7]. Of the 427 functions present before surgery, 26 were lost — 6% [7]. The mean number of performable functions rose from 4.2 to 9.3, and improved in 96 of 102 shoulders (94%); six shoulders had a net decrease, one of them losing six functions [7].
Set that cohort beside the normal shoulders of the same age and the whole argument fits in one picture.
Figure 3. The same twelve functions in three states: a normal shoulder aged 60 to 70 [28], an arthritic shoulder before total shoulder arthroplasty, and the same shoulder after [7]. Postoperative percentages are derived from the transition counts in Table 2. These are two different cohorts reported eight years apart, not one group followed through; the normal series does not report working full time. Sleep runs from universal, to nine in a hundred, to nearly nine in ten.
Now read the bottom row of Table 2. Fifty patients could not toss a softball underhand before surgery, and one of them regained it. It is the only one of the twelve functions that did not improve. The composite for this cohort went from 4.2 to 9.3 and rose in 94% of shoulders, and inside that result one function did not move at all.
The companion paper makes the same point from a different direction. In 124 shoulders with primary degenerative joint disease followed for a mean of 3.5 years, the number of performable functions rose from 3.8 to 9.3, and eleven of the twelve improved significantly [8]. The exception there was carrying twenty pounds at the side.
Put the two exceptions side by side and a pattern appears. In the 2002 series, 52 of 102 patients could already toss underhand before surgery and 68 of 102 could already carry twenty pounds — the two highest preoperative prevalences of the twelve [7]. The functions that fail to improve are the ones most patients had not lost. A change score cannot show this, because a function nobody lost contributes nothing to a change in either direction.
The 2002 discussion then provides the answers to the three questions patients ask: how much better will my shoulder be, in what ways will it be improved, and what are the chances I lose what I still have [7]. Roughly two-thirds of what was lost comes back. The chance of regaining a specified lost function is about 70%. The chance of losing a specified function the patient has before surgery is about 6% [7].
That is information a patient can use. It was published in 2002, and the transition table has not been repeated for this operation since. Figure 1-8 of the 1994 chapter had shown the same thing prospectively, function by function and interval by interval, eight years earlier [28].
The function-level question, asked again in 2026
The assessment of individual functions did return, twenty-four years later, from a different direction. In 1,048 patients with a minimum of two years of follow-up — 468 anatomic total shoulders, 164 reverse, 416 ream-and-run — each of the twelve Simple Shoulder Test functions was tested against reported satisfaction by receiver operating characteristic analysis [26]. Overall satisfaction at two years was 83.5% for anatomic total shoulder, 80.5% for ream-and-run and 68.9% for reverse [26].
Simple Shoulder Test item | aTSA (n = 468) | rTSA (n = 164) | RnR (n = 416) |
1. Comfortable with arm at rest | 0.221 | 0.409 | 0.400 |
2. Sleep comfortably | 0.360 | 0.548 | 0.675 |
3. Reach small of back | 0.239 | 0.036 | 0.325 |
4. Hand behind head | 0.299 | 0.470 | 0.450 |
5. Place coin on shelf | 0.377 | 0.367 | 0.175 |
6. Lift 1 lb to shoulder | 0.169 | 0.436 | 0.550 |
7. Lift 8 lb to top of head | 0.279 | 0.321 | 0.475 |
8. Carry 20 lb at side | 0.326 | 0.430 | 0.325 |
9. Toss softball underhand | 0.322 | 0.218 | 0.650 |
10. Throw softball overhand | 0.284 | 0.188 | 0.475 |
11. Wash opposite shoulder | 0.244 | 0.318 | 0.250 |
12. Work full time | 0.352 | 0.418 | 0.350 |
Table 3. Youden index for each Simple Shoulder Test function against two-year patient satisfaction, by arthroplasty type, 1,048 patients [26]. Higher values indicate an item that better separates satisfied from unsatisfied patients.
Sleeping comfortably had the highest Youden index of the twelve after reverse arthroplasty and after ream-and-run, and the second highest after anatomic total shoulder [26]. Comfort with the arm at rest carried a positive predictive value near 0.95 in all three groups, and in the reverse cohort its negative predictive value was 1.000 — every reverse patient who did not gain a comfortable arm at rest was unsatisfied [26].
The items that separate a satisfied patient from an unsatisfied one are, across all three operations, the first two on the list: comfort at rest and comfort at night. They are also the two functions that 80 of 80 normal shoulders could perform [28].
The rest of the table is a statement about who chooses each operation. After anatomic total shoulder the next best items are placing a coin on a shelf and working full time; after reverse arthroplasty, placing the hand behind the head; after ream-and-run, tossing a softball underhand [26]. Reaching the small of the back was the weakest item in the reverse cohort, with a positive predictive value of 0.543, the only value below 0.600 anywhere in the analysis [26]. And in all three groups, patients who did not regain overhand throwing were unlikely to report satisfaction, whatever else they regained [26].
7. Current use of the SST score
For anatomic total shoulder arthroplasty we set the threshold ourselves, and we set it lower than for any other procedure we perform. Across 887 arthroplasties, the anchor-based minimal clinically important difference for the Simple Shoulder Test was 2.3 overall, but 1.6 for anatomic total shoulder, against 2.6 for ream-and-run, 3.0 for hemiarthroplasty, 3.3 for cuff tear arthropathy arthroplasty and 3.7 for reverse [9]. Of the 368 anatomic patients, 96% cleared their threshold — the highest proportion of the five [9].
We then tested how well those thresholds identify a satisfied patient, in 406 anatomic total shoulders with two-year scores and a satisfaction anchor [10]. Sixty-four of the 406, 15.8%, reported satisfaction of mixed or worse [10].
Threshold | Value | Sensitivity | Specificity | Youden J |
MCID, anchor-based | 2.0 | 0.99 | 0.22 | 0.21 |
Substantial clinical benefit | 2.7 | 0.98 | 0.19 | 0.16 |
Change, ROC-optimized | 3.5 | 0.92 | 0.45 | 0.37 |
%MPI, anchor-based | 34% | 0.98 | 0.36 | 0.34 |
%MPI, ROC-optimized | 61% | 0.83 | 0.66 | 0.49 |
PASS, calculation-based | 10.2 | 0.55 | 0.81 | 0.36 |
PASS, 75th percentile | 12.0 | 0.33 | 0.88 | 0.21 |
Final score, ROC-optimized | 9.5 | 0.71 | 0.77 | 0.48 |
Table 4. Simple Shoulder Test thresholds tested against patient-reported satisfaction, 406 anatomic total shoulder arthroplasties [10]. %MPI, percentage of maximal possible improvement; PASS, patient acceptable symptom state. The source paper gives the calculation-based PASS as 10.2 in its threshold table and 11.0 in its predictor table; the value with sensitivity and specificity attached is used here.
Read the specificity column. At a minimal clinically important difference of 2.0, specificity is 0.22 — of the patients who told us they were dissatisfied, more than three in four had nevertheless cleared the threshold [10]. At the substantial clinical benefit of 2.7, specificity is 0.19 [10]. The measure we use to certify that an operation helped is satisfied by four of every five patients who say it did not.
And the choice of threshold changes not only who counts as a success but why. In those same 406 shoulders, a low preoperative score predicted success when success was defined as a change, while a high preoperative score predicted it when success was defined as a final score [10]. Same patients, opposite direction. Non–workers’ compensation insurance predicted success under every definition tested [10].
One cohort makes the size of the gap plain: among 188 anatomic total shoulders at a minimum of five years, 95% surpassed the minimal clinically important difference of 1.6, while only 62% reached an excellent final score of 10 or better [11].
Raising the threshold does not repair the disagreement. In a prospective series of 337 arthroplasties we defined a better outcome as at least 30% of the maximal possible improvement at two years with no second procedure — a considerably higher bar than any minimal clinically important difference [12]. Of the 237 patients who met it, 19 reported mixed feelings and 10 reported that they were mostly dissatisfied, unhappy or terrible. Of the 57 who failed to meet it, 15 were delighted, pleased or mostly satisfied [12]. The predictive model built from that definition identified the patients who would meet it with a sensitivity of 91% and those who would not with a specificity of 65%; in the simplified bedside version, specificity fell to 39% [12]. We are good at predicting the result we count as success and poor at predicting its absence.
The change and the final score are not two views of the same thing
An earlier study of ours shows this in the arithmetic itself. In 408 arthroplasties with two-year follow-up we set out to validate the Simple Shoulder Test against five clinical hypotheses [27]. The instrument did what was asked of it: scores rose from 3.9 to 10.2, with a Cohen’s d of 2.29 and a standardized response mean of 2.05, both larger than any SF-36 domain measured in the same patients [27]. What is worth reading now is which column reached significance.
Comparison | Preoperative | Final score | Change |
ASA 1 vs ASA 4 | 5.0 vs 1.0, P < .001 | 10.3 vs 5.3, P < .001 | P = .2 |
Narcotics before surgery vs none | 2.9 vs 3.9, P = .001 | 7.6 vs 9.4, P < .001 | P = .075 |
Previous failed surgery vs none | 3.4 vs 3.7, P = .260 | 8.2 vs 9.3, P = .002 | P = .049 |
Medicaid or workers’ compensation vs other insurance | 2.4 vs 3.8, P = .003 | 6.6 vs 9.3, P < .001 | 4.2 vs 5.5, P = .037 |
Table 5. Four of the five criterion-validity comparisons in 408 shoulder arthroplasties, by preoperative score, final score and change [27]. The ASA 4 group contained three patients.
In three of those four the final score separates groups that the change score does not, or barely does. A sicker patient, a patient taking narcotics before surgery, and a patient with a failed previous operation each end up lower than his counterpart, and each improves by about the same amount as everyone else. Counted as change, all three are ordinary successes. Counted as where they finished, they are not.
The satisfaction rows in that paper are the other half of it. Patients who called the result terrible had a mean final score of 2.7 and a mean change of −1.4; those who were delighted finished at 11.0 with a change of 6.9 [27]. Before surgery the two groups were indistinguishable, at 4.1 and 4.0 [27]. Where a patient started said nothing about whether he would be satisfied. Where he finished said almost everything.
How our studies compare with those of others
Everything above reflects our own practice, and a reader is entitled to ask whether the disagreement between threshold and satisfaction is peculiar to us. It is not.
A systematic review collected every minimal clinically important difference published for shoulder arthroplasty between 2008 and 2020: 43 studies, 16,408 patients, 17 different outcome measures, 112 thresholds [13]. For the ASES score the published values run from 6.3 to 29.5. For the Simple Shoulder Test they run from 1.4 to 4.0. For the Constant score they run from −0.3 to 12.8 [13]. Our 1.6 sits near the bottom of a range whose top is nearly three times as high.
That spread is not mostly a fact about patients. Of the 43 studies, 24 did not derive a threshold at all but took one from another paper, and 11 of those 24 took it from the same paper [13]. Of the 19 that did derive a new one, seven used an anchor; the review classifies three as arbitrary, defined there as using no known calculation technique, and two as not specifying a method [13].
The Constant value of −0.3 deserves its own sentence. A negative minimal clinically important difference means a patient can finish lower than he started and be counted as having gained a clinically important amount. The review reports a second one, −5.3 degrees for active external rotation after reverse arthroplasty, and both come from the same anchor-based calculation [13]. Its explanation is that the outcome measure and the anchor question are not tracking the same thing [13]. That is the specificity column of Table 4, reached from the other direction.
The largest threshold study in this field reaches the rest of the conclusion in its own words. In 5,851 arthroplasties performed by 38 surgeons at 30 sites, the minimal clinically important difference for the ASES score was 13.9, cleared by 91.3% of patients; for reverse arthroplasty performed for rotator cuff tear it was 4.4 points of 100, cleared by 94.7% [14]. Our 96% is not a local artifact. And it is worth stating how a number like that arises, because it is not what the name suggests: it is the difference between the mean improvement of the patients who called their shoulder better and the mean improvement of those who called it unchanged or worse, with everyone who called it much better excluded from the calculation [14]. Two group means, subtracted, then applied to individuals. Nothing in that arithmetic keeps the answer from being small, and nothing keeps it from being negative.
Across sub-cohorts defined by implant, diagnosis and sex, the patient acceptable symptomatic state varied by as much as 66%, and the minimal clinically important difference and substantial clinical benefit by as much as 1600% [14]. They conclude that these thresholds are fragile, that caution should be exercised in conflating them across studies, and that investigators should calculate the value in their own cohort rather than reference one from another paper [14]. The single most borrowed threshold in this literature is an earlier paper from that same group [13].
Why the change score behaves this way
There is an older result that explains all of it, published in the rheumatology literature in 2006 [15]. In 1,019 outpatients with knee osteoarthritis and 271 with acute rotator cuff syndrome, two quantities were derived from the patients themselves: the minimal clinically important improvement, taken from those who said their response to treatment had been good; and the patient acceptable symptom state, taken from those who said their current condition was satisfactory [15]. Then the arithmetic. In the lowest tertile of baseline pain, the mean starting score was 40.0 mm, the minimal important improvement was 10.8 mm, and the acceptable state was 27.0 mm [15]. Forty minus 10.8 is 29.2. The relationship held in every tertile, for pain and for function, in the chronic condition and the acute one [15].
The minimal important improvement is not an independent quantity. It is the distance from where the patient started to a state he would accept.
That has a direct consequence for us. A threshold derived in a population arriving at 3.6 of 12 [3] is a statement about where those patients started. Move the starting point and the threshold moves with it — which is what we found across our five procedures [9] and what the multicenter group found across implant, diagnosis and sex [14]. The variability is not a defect in anyone’s method. It is what the quantity is.
It also accounts for Table 5. If the threshold is baseline minus an acceptable state, then a change score is a final score with the baseline subtracted out — and the baseline is where the patient’s health, insurance and previous surgery had already left their mark. Subtract it and those groups become indistinguishable, which is exactly what the change column shows.
And what the number tracks is mostly the person
In 210 anatomic total shoulders at a mean of eight years, the final Simple Shoulder Test score, the change in score, and the percentage of maximal possible improvement were not correlated with humeral head centering, humeroscapular alignment, Walch classification, or glenoid version, and there were no preoperative radiographic predictors of a low final score [16]. Of 51 shoulders with radiographs at a minimum of five years, the 15 meeting criteria for radiographic loosening had a higher mean final score than those that did not, 10.3 against 8.7 [16]. That is not a single-series finding. Nine years earlier, in the prospective 337, neither preoperative glenoid version nor posterior decentering of the humeral head was associated with the two-year outcome [12].
Meanwhile, across 399 anatomic arthroplasties, resilience on the Connor-Davidson scale was independently associated with the Simple Shoulder Test, the ASES score and satisfaction [17]. In the five-year cohort, male sex predicted a final score of 10 or better with an odds ratio of 3.46 (95% CI 1.70–7.31), and workers’ compensation coverage predicted failing to reach it, 0.12 (0.02–0.60) [11]. In the 544 concurrent cases, a work-related shoulder problem cost 2.3 points on the final score (95% CI −3.5 to −1.1) [4].
We should be careful about what that shows. It shows that in this practice, with conservative reaming and no attempt at version correction, glenoid morphology did not predict the patient-reported result [12,16]. It does not show that glenoid morphology is irrelevant in general. What it does say is that our instruments are picking up the patient’s health, sex, circumstances and resilience at least as strongly as anything visible on the axillary view.
8. What we ask about the surgeon
Everything above concerns what we measure in the patient. There is one place where the field measures the surgeon instead, and it is worth looking at what that literature counts.
Surgeon volume has been examined in six large datasets, and it does affect the result. Higher volume is associated with fewer revisions, readmissions and early complications, at annual thresholds somewhere between 5 and 29 cases depending on the dataset [19,20,21,22,23,24].
Two features of that body of work bear on this post. The first is where the effect sits in time. The Australian registry finding is confined to the first 1.5 years [19] — the interval in which technical error declares itself. The claims data say the same thing from a different direction: reoperation runs at an odds ratio of 0.75 at 90 days, 0.86 at one year, and 0.94 and not significant at two [23]. The two-year figure contains the ninety-day one, so an effect confined to the early window does not reverse as follow-up lengthens; it dilutes. Two continents, two designs, the same shape.
The second is which outcome the whole literature uses.
Dataset | Volume contrast | Endpoints reported | Patient-reported outcome |
Australian registry, 2004–17 [19] | <10/yr vs >20/yr | Revision (HR 1.36, first 1.5 yr only) | None |
England, NJR + HES, 39,281 procedures [20] | Threshold 10.4/yr; also within-surgeon deviation | Revision, reoperation ≤12 mo, serious adverse events, prolonged stay | None (Oxford Shoulder Score linked and available) |
US national analysis [21] | ≥10/yr and ≥29/yr vs <4/yr | Revision within 2 yr | None |
Medicare, 90,318 reverse arthroplasties [22] | <28/yr and >96/yr, each vs medium | Readmission, transfusion, cost, length of stay | None (stated as a limitation) |
National claims, 155,560 arthroplasties [23] | Top decile vs the rest | Complications, readmission, reoperation at 90 d, 1 yr, 2 yr | None |
Meta-analysis, 332,542 patients [24] | 5/yr and 25–28/yr thresholds | Complication, revision (OR 1.41 each) | None |
Systematic review, 8 studies [25] | Any annual threshold | Length of stay, cost | Predefined as the primary outcome; not reported by any included study |
Table 6. What the surgeon-volume literature counts. The Medicare study’s reference group is the medium-volume group, not the highest, so a low-volume contrast there is a contrast against 28 to 96 cases a year. The claims study’s abstract and results give different surgeon counts; the results figures are used here, as they reconcile with that paper’s own fellowship table.
Not one of these analyses reports the proportion of patients regaining a named function, or reaching a score they would accept, or saying they were satisfied, by surgeon volume. The English study is the sharpest example: it drew on the one registry in the world with linked patient-reported outcomes, and the Oxford Shoulder Score was not among its endpoints [20]. The Medicare authors state the limitation themselves [22].
That absence is not something we are inferring. A systematic review set out to define an annual volume threshold and predefined its primary outcome as any patient-reported outcome. Eight studies met inclusion. The authors report that functional outcomes were not reported, and conclude that there is insufficient evidence that annual volume produces better patient-reported and functional outcomes — the published thresholds rest on length of stay and cost [25]. Someone went looking for exactly this and found nothing to pool.
What the effect is made of has been partly measured, and the answer bears on how a volume minimum should be read. In the English data, volume was entered twice: once as each surgeon’s mean annual volume across his whole record, and once as the deviation of any given year from that surgeon’s own mean. Mean volume predicted revision, reoperation, serious adverse events and length of stay. Deviation volume predicted none of them [20]. The deviation term is a within-person comparison — same training, same hands, same unit, same referral filter — and the only thing that varies is how many he did that year.
Whether these surgeons are busy because they are good, or good because they are busy, the design cannot say. But it can say that being busier this year than last did nothing.
Two of the candidate explanations have since been measured. Case mix is one: low-volume surgeons operate on sicker people, with a Charlson comorbidity index of 2.01 against 1.85 and more of every comorbidity examined [23], and on harder indications — proximal humerus fracture was the diagnosis in 15.8% of low-volume Medicare cases against 6.2% of high-volume ones [22]. Adjusting for all of that did not make the effect go away [22,23].
Selection is the other, and it has been measured in a form that fixes its direction. Among high-volume surgeons, 29.3% had completed a shoulder and elbow fellowship, against 12.1% of the low-volume group. Sports medicine fellowship did not differ at all, 39.1% against 40.6% [23]. The difference is not fellowship; it is this fellowship. And a fellowship is finished before a surgeon’s practice volume accrues, so volume cannot have produced it. Whatever part of the volume effect travels with shoulder and elbow training is not a dose of operating; it is something the surgeon brought to his first case. One finding cuts the other way and should be stated: the Medicare authors tested for an interaction between volume and fellowship and found none [22], so the training reading is supported by composition and not, so far, by effect.
That changes what a volume minimum would be doing. On the measured part of this it is partly a credential test and partly a statement about who was referred where. Both of those a patient can simply ask about — what fellowship did you do, and how many shoulders like mine do you see. Neither requires a hazard ratio. What no one can hand him is the thing he came in for.
This also constrains how any single-surgeon series should be read, including every one in this post. A surgeon operating well above all of these thresholds is not the surgeon the population figures describe, and a result obtained in that practice is a statement about what is achievable rather than about what is typical. It constrains this practice in a second way as well, and the case-mix finding is the reason. If low-volume surgeons are sent sicker patients, then a high-volume practice is receiving healthier ones. The 931 patients in section 2 and the 406 in section 7 were not a random sample of the shoulders that need arthroplasty. They were the shoulders that were sent here.
Part Three — Getting the question back
9. Name the target before the operation
Dr. C, in the 1993 worked example, is asked for his own results, function by function, for one operation and one diagnosis [18]. That is a smaller and more answerable question than any threshold in section 8, and it is the one the volume literature has never posed. Asking it requires no new instrument, no new software, and no new visit.
The preoperative Simple Shoulder Test is already in front of us when we obtain consent. Go through it with the patient and ask two questions. Of the items you cannot now do, which two or three do you most want back? And of the items you can still do, which would you least want to lose?
Record the answers before surgery. At each follow-up, report against them.
● Success is defined by the patient, in his own words, in advance — so it cannot be adjusted afterward to fit the result.
● It is auditable one patient at a time, which is what an individual surgeon needs and what no population threshold provides.
● It makes the preoperative conversation concrete, and we already have the numbers to answer it — Table 2 is that conversation.
● It captures loss, which no composite and no change score can.
Barron asked for something close to this in the same list of guidelines. The ninth reads that the assessment of the results of surgery should be made and agreed by the patient, the physician, and the surgeon [1]. He wanted the patient to be a party to the verdict. Fifty-seven years later the verdict is a threshold we set.
Barron was not the last to ask. In 2006 Tubach and colleagues concluded from the two cohorts above that what matters to a patient is feeling good rather than feeling better, and that in daily practice patients should be asked whether they feel good rather than whether they feel better [15]. Twenty years later we still report the change.
Their data also carry the warning that goes with this proposal. In the chronic condition, what patients called a satisfactory level of functional impairment depended on how impaired they had been: the least impaired called 20.4 satisfactory and the most impaired called 43.1, and the authors read this as patients becoming resigned to limitation and adapting to it [15]. Pain behaved differently. The acceptable level was stable whatever the starting point, and their reading is that one does not get used to a high level of pain [15].
Our patients arrive at 3.6 of 12 [3] after years of accommodation. If we ask only whether the result is acceptable, we are measuring against expectations the disease itself has lowered. That is an argument for asking before the operation, when the patient names the function he wants back, rather than after, when he is asked to rate what he was given.
10. Why counting loss matters most
A net change conceals what moved in each direction. A patient who goes from 6 performable items to 9 is recorded as plus three. He may have gained five and lost two, and the two he lost may be the two he cared about.
We have published the patients in whom that happened, and they are ordinary total shoulders for osteoarthritis. In the 124-shoulder series, five patients could perform fewer functions after surgery than before and three could perform the same number [8]. Those eight had better function than the cohort before surgery — a mean of 6.8 functions against 3.8 — and 5.3 afterward [8]. None of them had revision surgery [8]. In the 102-shoulder series, six shoulders had a net decrease and one lost six functions [7].
Those patients appear in no revision count. They appear in no survivorship curve. They appear in none of the six datasets in section 8, every one of which would have recorded them as uncomplicated successes. They are visible only because the preoperative answers were recorded item by item, and the authors of the 2001 paper say so directly: the lack of benefit becomes apparent only when both the outcome and the preoperative values are known [8].
One qualification. That series excluded patients covered by workers’ compensation, so it cannot be read alongside the workers’ compensation findings in section 7.
These eight are not errors in the data. They are the reason for the proposal.
11. What a paper written this way would report
● The named targets, collected before surgery and summarized across the cohort — which would tell the field what patients are actually asking arthroplasty to do.
● The proportion who regained the specific items they named. This is the primary outcome.
● The proportion who lost an item they named as one they did not want to lose, reported separately and never netted against gains.
● A within-person transition table for all twelve items, in the form of Table 2 — which we have known how to build since 2002.
● The composite score as well, so the work stays comparable to everything already published.
One denominator, stated once, used for all five.
12. What it would cost, and what it would settle
The cost is a question added to a form we already administer and a conversation we should already be having.
The gain is that several arguments dissolve rather than needing to be won. There is no dispute about which threshold to use, because the patient sets it. There is no puzzle about a patient who clears the minimal clinically important difference and is dissatisfied, because we can ask him whether he got back what he named. And a patient who is revised has by definition lost the functions he named, so the revision is carried in the same table as everything else.
It would also give the volume question something to measure. A threshold expressed in named functions regained is a threshold a surgeon can be held to and a patient can understand, and it does not require settling why busy surgeons do better before it can be reported.
One difficulty does not dissolve. Patients who lost the item they cared about are the least likely to answer the next survey, and no analysis recovers a person who has stopped replying. That is a problem of follow-up rather than of measurement.
A second caution: we do not yet know whether a named-target outcome would behave well as a research endpoint. It would need to be prespecified and standardized to be analyzable, and standardizing it may cost some of what makes the clinic conversation valuable.
A closing question
In 1987 we reported that 4 of the 50 shoulders in our first series could sleep on the affected side before surgery, and 43 of 50 could afterward [5]. In 2002, of 102 shoulders operated for osteoarthritis, 9 could sleep comfortably before surgery; 87% of the other 93 regained it, and none of the 9 lost it [7]. In 2018, across 931 patients arriving for arthroplasty, 9.6% could still sleep comfortably [3]. In 2026, across 1,048 patients, sleeping comfortably was the single function that best separated a satisfied patient from an unsatisfied one [26].
Four decades, four cohorts, one finding: the thing people lose is sleep, and the thing they come in for is sleep. We have known how to report that since the beginning, and we have known how to report the risk of losing what they still have since 2002.
What would we learn if the next hundred patients told us, before surgery, what would count as success — and we reported how often we delivered it?
Keeping eye on the most important thing
Peregrine Falcon
Skagit County
Follow on twitter/X: https://x.com/RickMatsen
Follow on facebook: https://www.facebook.com/shoulder.arthritis
Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133//
References
[1] Barron JN. Assessment of suitability for surgery in general. Timing of operation. Ann Rheum Dis 1969;28(Suppl 5):74-76.
[2] Matsen FA 3rd, Tang A, Russ SM, Hsu JE. Relationship between patient-reported assessment of shoulder function and objective range-of-motion measurements. J Bone Joint Surg Am 2017;99:417-426. doi:10.2106/JBJS.16.00556
[3] Somerson JS, Hsu JE, Neradilek MB, Matsen FA 3rd. The “tipping point” for 931 elective shoulder arthroplasties. J Shoulder Elbow Surg 2018;27:1614-1621. doi:10.1016/j.jse.2018.03.008
[4] Matsen FA 3rd, Whitson A, Jackins SE, Neradilek MB, Warme WJ, Hsu JE. Ream and run and total shoulder: patient and shoulder characteristics in five hundred forty-four concurrent cases. Int Orthop 2019;43:2105-2115. doi:10.1007/s00264-019-04352-8
[5] Barrett WP, Franklin JL, Jackins SE, Wyss CR, Matsen FA 3rd. Total shoulder arthroplasty. J Bone Joint Surg Am 1987;69:865-872.
[6] Richards RR, An KN, Bigliani LU, Friedman RJ, Gartsman GM, Gristina AG, Iannotti JP, Mow VC, Sidles JA, Zuckerman JD; Research Committee, American Shoulder and Elbow Surgeons. A standardized method for the assessment of shoulder function. J Shoulder Elbow Surg 1994;3:347-352.
[7] Fehringer EV, Kopjar B, Boorman RS, Churchill RS, Smith KL, Matsen FA 3rd. Characterizing the functional improvement after total shoulder arthroplasty for osteoarthritis. J Bone Joint Surg Am 2002;84:1349-1353.
[8] Goldberg BA, Smith K, Jackins S, Campbell B, Matsen FA 3rd. The magnitude and durability of functional improvement after total shoulder arthroplasty for degenerative joint disease. J Shoulder Elbow Surg 2001;10:464-469. doi:10.1067/mse.2001.117122
[9] McLaughlin RJ, Whitson AJ, Panebianco A, Warme WJ, Matsen FA 3rd, Hsu JE. The minimal clinically important differences of the Simple Shoulder Test are different for different arthroplasty types. J Shoulder Elbow Surg 2022;31:1640-1646. doi:10.1016/j.jse.2022.02.010
[10] Quinlan NJ, Dasari SP, Sharareh B, Levins JG, Whitson AJ, Matsen FA 3rd, Hsu JE. Do we need to reconsider how we gauge success after anatomic total shoulder arthroplasty? A study of thresholds optimized for patient satisfaction using the Simple Shoulder Test. J Shoulder Elbow Surg 2025;34:e694-e701. doi:10.1016/j.jse.2024.11.013
[11] Mills ZD, Schiffman CJ, Sharareh B, Whitson AJ, Matsen FA 3rd, Hsu JE. Anatomic total shoulder: predictors of excellent outcomes at five years after arthroplasty. Int Orthop 2024;48:1277-1283. doi:10.1007/s00264-024-06148-x
[12] Matsen FA 3rd, Russ SM, Vu PT, Hsu JE, Lucas RM, Comstock BA. What factors are predictive of patient-reported outcomes? A prospective study of 337 shoulder arthroplasties. Clin Orthop Relat Res 2016;474:2496-2510. doi:10.1007/s11999-016-4990-1
[13] Kolin DA, Moverman MA, Pagani NR, Puzzitiello RN, Dubin J, Menendez ME, Jawa A, Kirsch JM. Substantial inconsistency and variability exists among minimum clinically important differences for shoulder arthroplasty outcomes: a systematic review. Clin Orthop Relat Res 2022;480:1371-1383. doi:10.1097/CORR.0000000000002164
[14] Simovitch RW, Elwell J, Colasanti CA, Hao KA, Friedman RJ, Flurin PH, Wright TW, Schoch BS, Roche CP, Zuckerman JD. Stratification of the minimal clinically important difference, substantial clinical benefit, and patient acceptable symptomatic state after total shoulder arthroplasty by implant type, preoperative diagnosis, and sex. J Shoulder Elbow Surg 2024;33:e492-e506. doi:10.1016/j.jse.2024.01.040
[15] Tubach F, Dougados M, Falissard B, Baron G, Logeart I, Ravaud P. Feeling good rather than feeling better matters more to patients. Arthritis Rheum 2006;55:526-530. doi:10.1002/art.22110
[16] Sheth MM, Mills ZD, Dasari SP, Whitson AJ, Matsen FA 3rd, Hsu JE. Anatomic total shoulder arthroplasty for posteriorly eccentric and concentric osteoarthritis: a comparison at a minimum 5-year follow-up. J Shoulder Elbow Surg 2025;34:473-483. doi:10.1016/j.jse.2024.04.026
[17] Levins JG, Dasari SP, Quinlan NJ, Whitson AJ, Matsen FA 3rd, Hsu JE. Anatomic shoulder arthroplasty: the correlation between patient resilience, mental health, and outcome. J Shoulder Elbow Surg 2024;33:S9-S15. doi:10.1016/j.jse.2024.03.008
[18] Lippitt SB, Harryman DT 2nd, Matsen FA 3rd. A practical tool for evaluating function: the Simple Shoulder Test. In: Matsen FA 3rd, Fu FH, Hawkins RJ, editors. The shoulder: a balance of mobility and stability. Rosemont, IL: American Academy of Orthopaedic Surgeons; 1993:501-518.
[19] Brown JS, Gordon RJ, Peng Y, Hatton A, Page RS, Macgroarty KA. Lower operating volume in shoulder arthroplasty is associated with increased revision rates in the early postoperative period: long-term analysis from the Australian Orthopaedic Association National Joint Replacement Registry. J Shoulder Elbow Surg 2020;29:1104-1114. doi:10.1016/j.jse.2019.10.026
[20] Valsamis EM, Collins GS, Pinedo-Villanueva R, Whitehouse MR, Rangan A, Sayers A, Rees JL. Association between surgeon volume and patient outcomes after elective shoulder replacement surgery using data from the National Joint Registry and Hospital Episode Statistics for England: population based cohort study. BMJ 2023;381:e075355. doi:10.1136/bmj-2023-075355
[21] Best MJ, Fedorka CJ, Haas DA, Zhang X, Khan AZ, Armstrong AD, Abboud JA, Jawa A, O’Donnell EA, Belniak RM, Simon JE, Wagner ER, Malik M, Gottschalk MB, Updegrove GF, Warner JJP, Srikumaran U; Avant-garde Health and Codman Shoulder Society Value Based Care Group. Higher surgeon volume is associated with a lower rate of subsequent revision procedures after total shoulder arthroplasty: a national analysis. Clin Orthop Relat Res 2023;481:1572-1580. doi:10.1097/CORR.0000000000002605
[22] Girdler SJ, Maza N, Lieber AM, Vervaecke A, Kodali H, Zubizarreta N, Poeran J, Cagle PJ, Galatz LM. Impact of surgeon case volume on outcomes after reverse total shoulder arthroplasty. J Am Acad Orthop Surg 2023;31:1228-1235. doi:10.5435/JAAOS-D-23-00181
[23] Harkin W, Berreta RS, Williams T, Turkmani A, Scanaliato JP, McCormick JR, Klifto CS, Nicholson GP, Garrigues GE. The effect of surgeon volume on complications after total shoulder arthroplasty: a nationwide assessment. J Shoulder Elbow Surg 2025;34:1112-1119. doi:10.1016/j.jse.2024.07.025
[24] Daher M, Ilyas MH, Gonzalez-Morgado D, Steinmann SP, Abboud JA, Kassam HF. Defining surgeon experience thresholds for the reduction in complications and revisions after shoulder arthroplasty: a meta-analysis of 332,542 patients. J Shoulder Elbow Arthroplasty 2026;10(1-2). [page range to confirm at final proof]
[25] Kooistra BW, Flipsen M, van den Bekerom MPJ, van Raay JJAM, Gosens T, van Deurzen DFP. Shoulder arthroplasty volume standards: the more the better? Arch Orthop Trauma Surg 2019;139:15-23. doi:10.1007/s00402-018-3033-7
[26] Collins AP, Gregory J, Whitson AJ, Matsen FA 3rd, Hsu JE, Schiffman CJ. Which shoulder functions correlate with patient satisfaction after primary shoulder arthroplasty? J Shoulder Elbow Surg 2026;35:19-27. doi:10.1016/j.jse.2025.04.016
[27] Hsu JE, Russ SM, Somerson JS, Tang A, Warme WJ, Matsen FA 3rd. Is the Simple Shoulder Test a valid outcome instrument for shoulder arthroplasty? J Shoulder Elbow Surg 2017;26:1693-1700. doi:10.1016/j.jse.2017.03.029
[28] Matsen FA 3rd, Lippitt SB, Sidles JA, Harryman DT 2nd. Practical evaluation and management of the shoulder. Philadelphia: W.B. Saunders; 1994. Chapter 1, Evaluating the shoulder; p. 1-17.
Disclosure: I have no financial relationship with any orthopaedic device manufacturer. Every patient series from our own institution cited above includes me as an author, as do references [18] and [28]. References [1] and [6] are not case series; references [13], [14] and [15] are a systematic review, a multicenter registry analysis and a rheumatology cohort from other groups; and references [19] through [25] are national datasets and reviews from other groups. Reference [14] draws on a registry funded by an implant manufacturer, and most of its authors have financial relationships with that manufacturer; the conclusion cited here runs against that interest. The criticisms in Part Two are criticisms of our own work first, and section 8 is a criticism of a literature in which I have not published. Figures 2 and 3 are original, drawn from published data in references [6], [7] and [28].



