Saturday, July 25, 2026

Periprosthetic infections in the JBJS - PJI of the shoulder is its own thing



One stage or two stage for PJI? For the shoulder, it's a lot more complicated than that


A perspective to start our discussion

A practical definition of bacterial infection is "bacteria doing harm". It is estimated that the healthy human harbors 38 trillion bacteria, a number essentially equal to the number of human cells in the body [27]. It is a wonder that all these bacteria in our shared ecosystem rarely do us harm. There is a peaceful coexistence — we and our bacteria keep each other healthy unless something disturbs the equilibrium. Escherichia coli lives quietly in the gut until something tips the balance and it causes colitis. Cutibacterium lives quietly in and around the shoulder, keeping the skin healthy until something tips the balance and harm results.

In shoulder PJI, the balance has been disrupted: the host is being outmatched by the bug. The goal of treatment is to restore the balance. The keys to accomplishing balance may include lowering the bacterial load, removing biofilm-containing implants, restoring stability of the articulation, and enabling the soft tissues surrounding the joint to become healthy, well vascularized, and functional in spite of the inevitable presence of ambient bacteria.

This concept is consistent with what we observe. For example, single-stage revision of a failed arthroplasty frequently yields durable improvement in patient comfort and function, even when deep specimens taken at the time of the revision return positive for Cutibacterium [7].

One stage or two stage for hip and knee PJI 
— does it matter?

A recent trial in the Journal of Bone and Joint Surgery compared one-stage with two-stage revision for chronic periprosthetic joint infection (PJI) of the hip and knee. Three hundred and twenty-three patients were randomized. At two years the success rate was reported as 97% after one stage and 91% after two; one stage was statistically noninferior to two [13,14]. The investigators deliberately included the patients that earlier single-stage series had excluded: those with draining sinuses, comorbidity, and antibiotic-resistant organisms [13]. But eligibility still required a chronic infection with a known organism; culture-negative infections were excluded, as were patients with fungal infections, immunosuppression, prior revision, and soft-tissue involvement that precluded wound closure [13,14].

"Success" was determined by a composite of several factors: no clinical failure or reinfection with the same or a new organism, no reoperation for PJI, and no PJI-related death [13]. That composite counts as a success the patient whose infection is controlled but who remains on suppressive antibiotics — 8.1% of the one-stage group and 17.1% of the two-stage group at final follow-up [13]. Infection control off antibiotics altogether, which is closer to what we would call a good result, was reached by 85.9% of the one-stage group and 70.7% of the two-stage group [13]. Note that success was not defined as the eradication of bacteria from the joint.

Of the 323 patients randomized, 258 had two-year follow-up, 16 patients in the one-stage group and 9 in the two-stage group having died before that point [13]. Nine patients in the two-stage group still had their spacers at two years; they were excluded from the analysis rather than counted as failures [13]. Not everyone assigned to two stages receives the second one.

The two arms also differed systematically in at least one important way beyond the number of operations: the arm that appeared to do better received an antibiofilm adjunct that the other arm almost never did. Rifampin was given to 29 of 84 patients with a staphylococcal organism in the one-stage arm, against 1 of 90 in the two-stage arm [13]. This practice follows infectious disease guidelines, which tie the use of adjunctive rifampin to debridement with implant retention or to one-stage revision [13] — but it is a difference in treatment all the same. The debridement and irrigation protocols were the same in both arms [13,14]. The antibiotics matched in kind but not in schedule; the two-stage arm carried the longer course, six weeks of intravenous therapy after the resection and at least six months of oral therapy after reimplantation, and did not do better for it [13]. Whether the one-stage result reflects the single operation or the drug that accompanied it is not a question this trial can settle.

The authors noted the success rate was above 90% in both arms, better than either strategy previously reported in the published literature; they attribute this to a protocol-driven treatment algorithm rather than to the choice between one stage and two [13]. Their conclusion: one-stage treatment should be strongly considered as the standard of care provided the indications and protocols described are explicitly followed [13]. In agreement, the commentary concludes that two-stage exchange should not be the de facto gold standard for chronic PJI in all patients with hip and knee PJI [14].

The shoulder is different

In the hip and knee trial, one stage versus two was a decision about a joint that is obviously infected by an organism already identified — most often a staphylococcus [13].

In contrast, while some shoulder periprosthetic infections are obvious (fever, swelling, pain, draining sinus, pus, positive cultures for Staphylococcus or a Gram-negative organism on joint aspiration), when we are doing a revision for a clinically failed shoulder arthroplasty we are most often considering a joint that may or may not be infected, by an organism that may never be identified at all.

The organism most commonly isolated at revision shoulder arthroplasty is Cutibacterium [2,23], a normal inhabitant of the pilosebaceous glands abundant in the healthy skin overlying the shoulder [4]. It is often indolent, low in virulence, and slow to grow. Cutibacterium PJI can present as the delayed, otherwise unexplained onset of pain and stiffness after a honeymoon period of usual post-arthroplasty recovery, without swelling, tenderness, abnormal blood tests, or positive cultures of a joint aspirate [1,2] — a "stealth" rather than an "obvious" presentation. Substantial cultures of this organism can be found in shoulders that appear entirely aseptic and are revised three years or more after the index arthroplasty [6]. The results of cultures taken at surgery are not known until weeks after the patient has left the operating room, so they cannot inform the decisions we make during the case.

For these reasons, we lean toward treating most failed shoulder arthroplasties as possibly infected. We are treating the individual patient, not just the shoulder and not the culture report. Our goal is to choose the procedure that offers the patient the best chance of functional recovery, that minimizes the risks we impose in the process, and that helps them restore balance in their relationship with the organisms in their shoulder's environment — recognizing that we will have no way to know if all the bacteria have been cleared from the shoulder [22].

The decision to perform a prosthesis exchange rests on the clinical judgment of the surgeon, weighing the concern for infection against the risks to the patient of implant removal and replacement. When the clinical picture is concerning for a shoulder PJI, we usually consider a thorough debridement with single-stage revision followed by empiric oral antibiotics until the results of a standardized set of cultures are finalized. We interpret those cultures by considering the bacterial load across multiple specimens [1] rather than by looking only at the number of specimens that are culture positive.

Because of the increased risk of a two-stage approach, we consider it primarily in severe obvious infections, sepsis, or failed single-stage revisions in patients who are physically, medically, immunologically, nutritionally and emotionally optimized for a long course of treatment and a second surgery. Following the same logic, we rely primarily on oral antibiotics because of the adverse outcomes related to intravenous antibiotic administration (line infection, thrombosis, emboli) and high-dose antibiotics (gut, kidney, liver, and nerve complications).

The problems with cultures

As surgeons we would like to know at the time of a revision procedure whether the shoulder is infected and, if so, by what bug. Cultures provide the most definitive evidence of PJI, but not in a timely manner, so we must determine what surgery to do and what antibiotics to use in the absence of the key information upon which those decisions would ideally be made.

Bacteria are not evenly distributed in the infected shoulder; this creates a sampling challenge. In shoulders with at least one positive culture, roughly half of the individual specimens grow nothing (mean 43%, median 50%) [1]. An infected shoulder sampled only once or twice may be read as culture negative depending on where the rongeur happened to go. The same pattern appears in the larger series: among shoulders that cultured positive for Cutibacterium, an average of 2.4 of 4.3 specimens were positive while 1.9 were negative [2].

The yield also depends on the specimen: fluid is the weakest source at 32.6% positive, compared with 66.5% for soft tissue and 55.6% for explants [1]. Those were specimens taken at the time of surgical revision rather than office aspirates, but the reason — that this organism lives in biofilm rather than free in joint fluid — applies to both, so a negative aspirate carries little weight. Bacteria are usually sessile rather than planktonic.

In addition, the organism may be genetically mixed: in a series of eleven shoulders selected for substantial bacterial burden, five had more than one subtype of Cutibacterium in the deep tissues despite similar colony morphology, four of them with two subtypes and one with four [3]. A single deep specimen may therefore misrepresent what is present elsewhere in the same shoulder.

Four more variables need consideration.

(1) How long the plates are held: only 45% of Cutibacterium cultures had turned positive at one week, 86% at two weeks, 97% at three, and 100% at four [2]. A laboratory holding for only five days would call most positives as negative. Working the other direction, every infected event in a dedicated culture study had declared by day 13, while 21.7% of the nondiagnostic events did not turn positive until after day 13 [5]. Holding beyond two weeks may not be clinically useful.

(2) Which media are used: a diagnosis would have been missed in 29.4% of infected patients had extended incubation been applied only to the anaerobic media [5].

(3) How many specimens are submitted: the number that turn positive rises with the number cultured [2].

(4) Whether antibiotics were held until cultures were obtained, which may improve diagnostic accuracy [2,15]: positive cultures for Cutibacterium and for other organisms were each more than twice as likely when antibiotics had been withheld until specimens were harvested [2].

The case for standardizing specimen harvest, culturing and result interpretation

Because the specimen count, the media, and the hold time vary from case to case, "two positives" in one patient and "two positives" in another may not have the same importance. If we are to compare culture results among patients and among institutions, standardization is important. While there are various recommendations in the literature [1], our protocol is 5 deep tissue specimens, each obtained with sterile instruments and cultured in broth and aerobic and anaerobic media and observed for two weeks.

No number or percentage of positive specimens has been shown to mark a clinically useful threshold for infection [1]. The reason is apparent from the sampling challenge discussed above. The often used threshold of "two or more positive cultures" depends on how many samples were taken and from where as much as on the amount of bacteria in the shoulder.

The more informative finding may be the load — the degree of positivity across an adequate set of specimens — read together with the clinical picture. There are at least two ways to get a handle on the bacterial load in a shoulder.

(1) Report positive cultures as a fraction of the total submitted. This is the recommendation of the 2018 International Consensus Meeting [9] and of the 2014 review, which asked for the number of positive specimens divided by the number submitted [22]. This approach works best only if the number of specimens per shoulder is consistent. One out of two and three out of six both are 50%, but they may not have the same clinical significance.

(2) Examine the culture plates for the density of growth [1]. Standard plate streaking technique enables the laboratory to view growth in each of four quadrants, so the report can be 0, 1+, 2+, 3+, or 4+ indicating the number of quadrants with growth. Each result then carries a Specimen Propi Value — 0.1 for growth in broth only, 0.1 for a single colony on the plate, then 1, 2, 3, and 4 for the number of quadrants — and the sum of those values across a shoulder is its Shoulder Propi Score, which divided by the number of specimens submitted gives the Average Shoulder Propi Score [1,9]. Load defined this way separates the groups: specimens from infected patients were 6.3 times more likely to demonstrate growth on two or more of the four quadrants (p = 0.002) [5].

Environmental contamination

Periodically, control specimens (such as a sterile sponge open to OR air during the case) should be submitted to check for environmental contamination in our operating room and microbiology laboratory [23]. If a fraction of the "sterile" control specimens are culture positive, this result should influence the interpretation of the specimens from the patient. Background rates are usually not zero. In a prospective study of 117 open deltopectoral procedures performed in shoulders with no suspicion of infection, a square of sterile gauze was cut with sterile scissors as the trays were opened, placed in a sterile container without ever being touched by a gloved hand, and sent to the laboratory alongside the tissue specimens. Seven of the fifty-four control sponges — 13.0% — grew bacteria, five of them revealing Cutibacterium at a median of fourteen days [23]. Across the same series, 20.5% of the surgeries yielded at least one positive tissue specimen, 18.3% among the shoulders with no previous surgery, and the difference between the sponges and the tissue did not reach significance (p = 0.234) [23]. The authors read this as evidence that reported rates of positive cultures at primary and revision arthroplasty need to be considered in light of the local levels of environmental contamination [23].

What raises the odds of a positive culture?

Patient sex

Male sex carried an odds ratio of 14.1 for a positive tissue culture, with a confidence interval running from 4.9 to 40.0, in shoulders that were not infected [23]; male sex also raises the odds of a positive Cutibacterium culture roughly sixfold in shoulders revised for pain, stiffness, or loosening [2]. Taken together, these results indicate that male sex is associated with culture positivity, and leave open how much of that association is about infection rather than about how much Cutibacterium a man's shoulder carries into any operation.

Loosening can be a presentation of Cutibacterium PJI

It is tempting to consider a loose component as a mechanical problem rather than an infection. In the shoulder that assumption does not always hold. In 193 arthroplasties revised for pain, stiffness, or loosening — that is, without obvious infection — 56% had positive cultures, and the odds of a positive Cutibacterium culture were roughly threefold higher with humeral component loosening, fourfold each by glenoid wear and by membrane formation, sixfold by male sex, tenfold by humeral osteolysis, and twelvefold by cloudy joint fluid [2].

What we have learned about single-stage revision
 in the shoulder

When we revise a failed hemiarthroplasty or anatomic total shoulder in a single stage, patients often do well even when several deep specimens later return positive for Cutibacterium. In 55 revisions performed without clinically obvious infection, the 27 shoulders with two or more positive cultures at the time of revision improved their Simple Shoulder Test scores from 3.2 to 7.8 at a mean of just under four years, which was at least as good as the 28 control shoulders (2.6 to 6.1) [7]. Eleven percent in each group required a further procedure for persistent pain or stiffness [7].

This was not a perfect study. The control cohort was defined by having no more than one positive culture, and the two groups were not treated alike: patients whose cultures reached two positives received six weeks of intravenous antibiotics with oral rifampin, followed by at least six months of oral antibiotics, while the controls stopped at three weeks [7]. The groups also differed sharply in a variable known to predict culture positivity: 89% of the culture-positive patients were men, against 39% of the controls [7]. In spite of these limitations, what the study does show is that shoulders with two or more positive cultures at the time of revision improved after a single-stage exchange by about as much as shoulders with fewer. In other words, patients can improve even when a revision implant is placed in a contaminated field.

Reinfection after revision

A systematic review with meta-analysis found reinfection after single-stage revision no worse than after two-stage — 6.3% versus 10.1%, a difference that did not reach significance [8]. Its authors attribute the apparent single-stage advantage to treatment bias rather than to the operation: the single-stage cohorts contained more Cutibacterium (48.7% versus 33.7%) and more acute and subacute infection, while the two-stage cohorts contained more MRSA (9.7% versus 2.5%) and more chronic infection [8]. Reinfection was defined by each included author's own criteria, which the review concedes are highly variable [8]. Its conclusion was that a surgeon treating Cutibacterium or another sensitive, low-virulence organism with a single-stage exchange is likely to have a low recurrence rate [8].

The consensus meeting reached a similar conclusion. It pooled 161 single-stage and 325 two-stage shoulder revisions from the published literature and found reinfection in 5.6% after single-stage and 11.4% after two-stage, with complications in 12.7% and 21.9% [10]. Constant-Murley scores were similar: 49.1 and 51.1 [10]. The delegates pointed out that surgeons had plausibly routed the worse infections to two stages and the milder ones to single stage, and that this alone could account for the gap [8,10].

The same proceedings also report failure, an endpoint they never define, and they tabulate it twice from overlapping but different sets of series, with opposite results. Quoting an earlier systematic review, they report failure in 9.9% after single-stage exchange, 6.3% after two-stage exchange, 9.7% after explantation with a permanent spacer, and 31.4% when the implant was retained [10]. Their own table, restricted to subacute and chronic infection and adding three later series to that review, reports 8.2% after single-stage exchange (33 of 404), 11.2% after two-stage (24 of 214), and 31.3% with retention (26 of 83) [10]. The two disagree about which exchange strategy does better, and the pool behind them is weighted opposite to the pool behind the reinfection figures — 404 single-stage against 214 two-stage here, 161 against 325 there [10]. These pooled numbers cannot settle the choice between one stage and two. What both tabulations agree on is that retaining the implant is associated with three to four times more failures than exchanging it.

The downsides of a two stage

A second stage costs the patient: another anesthetic, another operation, spacers that fracture or dislocate, and cuff and bone stock lost along the way, with function suffering for it [10]. The spacer may help control the organism, but no shoulder series has isolated what the spacer itself contributes, and the pooled data do not show the staged approach buying a lower reinfection rate [8,10]. What has been counted is what it costs. In the largest dedicated series, 60 spacers were placed in 53 patients — 39 for infection at the site of a shoulder arthroplasty, the rest for non-arthroplasty or primary shoulder infection [25]. Of the 44 patients who went on to a second stage after a mean interval of six months, 14 had 18 complications: eight bone erosions, four fractures of the spacer, three rotations of the spacer within the humeral shaft, and three humeral fractures, two of which required reoperation [25]. The rate was lower in the shoulders infected at the site of an arthroplasty than in the others, 27.3% of spacers against 62.5% [25]. All four spacer fractures happened during removal, when the head separated from the stem; in one, cement down the shaft forced a diaphyseal osteotomy and produced a humeral fracture that had to be cabled [25]. Two of the three humeral fractures happened with the spacer in place and without any trauma — one at two weeks in thin bone, one a greater tuberosity at four months [25]. At the second stage, nine greater tuberosities fractured while the reverse was being implanted, which the authors attribute to extreme scarring rather than to the spacer itself [25].

A separate series speaks to what happens when the second stage never comes. Seventeen patients whose spacer was placed as definitive treatment, with no intention to convert, had a mortality rate of 52.9% at a mean of 1.8 years after placement; five of the seventeen required a spacer exchange for persistent infection; and the eight who survived had a mean ASES score of 33.9 and a mean SANE score of 35.6 at a mean of 4.7 years, with a trend toward lower scores in those with type 3 humeral bone loss [26]. These patients were selected for a permanent spacer because they were older and carried more comorbidity than those who went on to reimplantation [26], so the mortality reflects who they were rather than what the spacer did. What the function scores show is the cost of a first stage that is never followed by a second.

The cost is not only mechanical. A report from the hip and knee randomized trial compared the first stage of a planned two-stage exchange with a one-stage exchange — a design that isolates the spacer, since the treatment was otherwise the same — and found acute kidney injury, defined as a creatinine at least 1.5 times baseline or a rise of at least 0.3 mg/dL, in 22.7% of the two-stage patients against 6.6% of the one-stage patients (p = 0.011), which is 15 of 66 against 4 of 61; spacer placement carried an odds ratio of 7.48 (95% confidence limits 1.77 to 31.56) on multivariable analysis [24]. Those were hips and knees, and no shoulder series has measured this.

What we have learned about antibiotics after revision

Having chosen a single-stage revision, the surgeon must still decide what antibiotics to give while the cultures are incubating. Again, this decision is made without the information that would settle it.

In a series of 175 revision shoulder arthroplasties, the route was chosen by the surgeon: three weeks of intravenous antibiotics when the index of suspicion for infection was high, three weeks of oral antibiotics when it was low, with the regimen modified once the cultures returned [11]. Male sex, a history of infection, intraoperative membrane formation, and younger age independently predicted starting intravenously [11]. The surgeons' preoperative and intraoperative impression predicted the culture result in about three-quarters of cases; a quarter of patients had their therapy changed after the cultures came back [11].

Complications were less frequent among those treated orally and for a shorter course; the orally treated patients were also the low-suspicion patients. However, within the two groups whose antibiotics stopped at three weeks the complication rate was 23% for intravenous and 6% for oral (p = 0.039), which is 3 of 13 against 5 of 83 [11]. Baseline risk still differs between those groups, so this is not a clean comparison of routes; but duration alone does not account for the difference.

Adverse effects of antibiotic administration occurred in 19% of patients in that series [11], and 14 of the 33 patients who were asked (42%) in the earlier one reported side effects [7]. Among the 92 patients who received a PICC line, 4% developed an upper-extremity venous thromboembolism, 3% had catheter migration, 2% could not have the line placed, and 9% had symptomatic skin irritation [11]. Three of the four thromboembolisms occurred in patients with initial inpatient PICC lines whose cultures were later seen to be negative [11], and patients started intravenously whose cultures ultimately came back negative had a 36% rate of antibiotic-related complications [16]. PICC lines placed as outpatients — by protocol the lines in the group started orally and escalated when cultures turned positive — carried a 30% complication rate compared with 13% for the lines placed before discharge [11]. That escalated group had the highest rate of antibiotic-related complications of any group in either report: 12 of 30 (40%) in the first and 8 of 15 (53%) at mid-term follow-up [11,16].

Apparent infection-free survival was 91% [16]. Those who started orally and were converted to intravenous therapy after positive cultures did about as well as those who started intravenously [16]. These groups were assigned by surgeon suspicion rather than randomized, so they are not a head-to-head comparison of routes. What the follow-up shows is that a protocol beginning with oral antibiotics and escalating on the basis of culture results did not appear to cost these patients control of their infection. The median improvement in the Simple Shoulder Test was 3 points and the median reduction in pain 4 points, both exceeding the minimal clinically important difference for those instruments, with a final median SST of 7, ASES of 62, and SANE of 60 [16]. Seventeen of the 92 (18%) underwent a further revision for any cause, 8 of them with two or more positive cultures [16].

The broader orthopaedic evidence is consistent with an oral-first approach. The OVIVA trial found oral antibiotics noninferior to intravenous antibiotics for bone and joint infection [17] — a well-powered randomized result, but not a shoulder result. A multicenter series of 172 prosthetic joint infections reported treatment failure in 9.2% of patients given oral therapy and 15.6% of those given intravenous therapy [15,19]. That difference was not statistically significant (p = 0.211), and route did not emerge as a risk factor in the multivariable model; knee infection and polymicrobial infection did [19]. Hips and knees made up 151 of the 172 [19]; only 10 were shoulders, and only one of the 22 recurrences in the whole cohort was a shoulder [15]. Seventy-six percent of the cohort was managed with debridement and implant retention rather than exchange [19], so the population was weighted toward less aggressive infection. The oral arm also combined 40 patients given oral therapy alone with 36 given up to three days of intravenous therapy first, and carried more Cutibacterium than the parenteral arm (13% versus 3%, p = 0.029) and fewer streptococcal and enterococcal infections [19], so organism and route are confounded there too.

In another small study nine Cutibacterium prosthetic joint infections were treated with single-stage exchange and oral antibiotics — six of them shoulders. Long-term follow-up was obtained on eight, and seven of those eight reported no further symptoms of infection at three years [18]. Every patient received linezolid together with rifampin, for a median of twelve weeks and a range of eight to twenty-four, and the diagnosis in each case rested on a single intraoperative culture together with pain or swelling [18].

Rifampin appears both in the hip and knee trial and in that nine-patient series. In staphylococcal periprosthetic infection it is an established antibiofilm adjunct, and the guidelines tie its use to implant retention and to one-stage revision [13]. For Cutibacterium the position is considerably weaker. Synergy has been shown in the laboratory, but the consensus meeting judged the clinical experience insufficient to endorse its use and noted that the literature on combining it is conflicting [10]. In the one shoulder series that used it heavily — rifampin in 15 of the 21 patients who received antibiotics — it did not change the outcome: 73% favorable with it against 60% without, p = 0.61, which is 11 of 15 against 3 of 5 [21]. Six of those 15 patients, 40%, had to stop the drug for adverse reactions, ranging from gastrointestinal and influenza-like symptoms to angioedema and a rash requiring hospitalization [21]. In our experience rifampin interacts with many pain medications by speeding up how the liver breaks them down, which can cause poorer pain control.

Our opinion is that oral antibiotics are a reasonable choice when revising a shoulder arthroplasty without purulence, sinus tract, or a virulent organism, while intravenous therapy may be appropriate for the patient with obvious infection, a large biofilm burden, prior infection, or immunosuppression [15].

How do we know whether the infection has been "cured"?

We cannot know whether an infection has been eradicated. Patients who are improved are not reoperated [7], but whether the organism was cleared, whether it persists under control, or whether it was never the cause of the failure are three possibilities that a good clinical result does not distinguish among. Showing eradication would require five deep samples cultured in a deliberate manner, which can only be accomplished by a return to the operating room for another big surgery. The problem is even greater for culture-negative infections: if we never identified an organism, we have nothing to declare eradicated.

Which is why "cure" is the wrong target. We are not trying to sterilize a joint. We are trying to leave the patient with the best comfort and function we can, in balance with ambient bacteria, with the least risk.

An algorithm for the possibly infected shoulder

Somewhat embarrassingly, an operational definition of a "true" periprosthetic shoulder infection still eludes us [12]. While the 2018 International Consensus Meeting framework tries to place a shoulder on the spectrum of probability — unlikely, possible, probable, or definite [9] — this algorithm is difficult to use in determining treatment. For example, "definite PJI" by culture requires two positive tissue cultures with phenotypically identical virulent organisms, but as we have seen, the culture results are not available when the initial treatment decisions need to be made. Furthermore, the criteria for diagnosing infection from the most common organism, Cutibacterium, are even less helpful; a shoulder growing this organism cannot reach "definite" with any number of cultures. Two of five cultures positive for Cutibacterium with well-fixed components and everything else negative is a "possible PJI", a designation without impact on clinical decision making.

Instead of trying to categorize probabilities in the absence of the key data, we consider every failed arthroplasty as possibly infected. If revision surgery is performed, the goal is to restore comfort, function, and soft tissue health by the least risky approach.

Here's our approach
1. Before surgery

• Assume a failed arthroplasty may be infected. Raise that suspicion in the presence of recognized Cutibacterium markers: male sex, humeral loosening or osteolysis, glenoid wear, membrane formation, cloudy joint fluid, a history of infection, and the stealth presentation of pain and stiffness without obvious infection [2,6,11]. Recognize that male sex predicts a positive culture in shoulders that were never infected at least as strongly as it predicts one in shoulders that were [2,23].

• Recognize that a normal ESR or CRP does not exclude infection. Among the culture-positive patients in our largest series, only 17% had an elevated ESR, 13% an elevated CRP, and 9% an elevated white blood-cell count, and none of the three was significantly related to culture positivity [2]. Another series of shoulders revised with positive intraoperative cultures shows the same pattern from outside our institution: an elevated CRP in 25% of those tested and an elevated ESR in 14% [22]. A negative aspirate carries little weight either, fluid being the weakest specimen [1].

2. At surgery — sample well, and rebuild for the host

• Harvest at least five deep specimens from different sites — capsule, humeral canal, collar and periprosthetic membranes, explants — favoring soft tissue and explant over fluid [1,7,9]. Use a fresh, individually peel-packed sterile instrument for each specimen, opened just before sampling, avoiding contact with the dermal structures [3,4]. Submit them for extended broth, aerobic, and anaerobic culture held for 14 days [5,9].

• Consider sending a control specimen with the case, at least periodically — a square of sterile gauze opened, handled, transported, and cultured exactly as a tissue specimen is. A laboratory returning growth on 13% of control sponges affects how we interpret cultures from the patient's shoulder [23].

• Send tissue for frozen section histology, noting that only 40% of cases with two or more positive specimens showed acute inflammation [5], and in our own series acute inflammation, chronic inflammation, and foreign-body reaction were each unrelated to Cutibacterium culture positivity [2].

• After thorough debridement, determine the potential benefit and risk of prosthesis exchange, recognizing that exchange removes a potentially biofilm-laden component but can damage the bone in the process. The decision is easy, of course, when the implant is loose.

• Choose a reconstruction that is durable and biologically sound: a stable, well-fixed construct that improves the health of the site. In our published series this most often meant a single-stage conversion to a hemiarthroplasty with an antibiotic-soaked allograft, to a total shoulder where glenoid bone stock allowed, or to a reverse where the cuff would not support anything else [7].

• Cover the likely organisms with an oral agent — amoxicillin-clavulanate or doxycycline, the agents used in our protocol [11] — rather than placing a PICC line for empiric intravenous therapy. Intravenous antibiotics have not been shown to give a better result in this setting, and their adverse effects are significant [11,16].

• Reserve a two-stage approach for the shoulder that is clearly infected with a virulent organism, has gross purulence or a sinus tract, has a large soft-tissue deficit, or has already failed a single-stage attempt — the situations where the host needs more help than a single operation can give.

3. When the culture results become available

• Record each specimen not only as positive or negative, but as the bacterial load, using the Specimen Propi Values described above; their sum is the Shoulder Propi Score and, divided by the number of specimens submitted, the Average Shoulder Propi Score [1].

• Consider the number of specimens growing bacteria, the load in each, and — if the laboratory will report it — the number of component media positive for growth [1,5].

• Modify the antibiotic regimen in response to the culture results, but escalate with caution. A high load, or multiple concordant specimens, supports treating as infection with organism-specific, infectious-disease-directed antibiotics [11,15]. Weigh that against the finding that concordant positives in more than one specimen occurred in 8.5% of shoulders with no infection at all [23], against the consensus finding that antibiotics continued beyond twenty-four hours after unexpected positive cultures for an indolent organism did not appear to reduce subsequent infection [10], and against the adverse-effect rates set out above, in which the escalated group fared worst of all [7,11,16]. Escalate for the organism and for the clinical picture, not for the culture report by itself.

• An isolated low-load single positive is often a contaminant, but we want to mark that as a judgment rather than a finding — no study has established a threshold below which a positive culture can be dismissed [1]. A patient with only one of four periprosthetic specimens positive has gone on to recurrent Cutibacterium infection [20, as reported in 5], and recovery from a single culture medium does not exclude clinical infection [5].

• Do not read negative cultures as proof of absence, particularly when few specimens were taken or the hold was short.

4. Follow-up

• Follow for recurrence of symptoms over time [12].

• Judge success by whether the patient has a stable implant and durable, comfortable function.

Conclusion

So one stage or two stages may be the wrong question for the patient with a failed shoulder arthroplasty. The hip and knee trial found an answer in a setting we shoulder surgeons rarely have — a known organism, a joint we already agree is infected [13].

If revision of a failed shoulder arthroplasty is planned, the available experience points toward a single operation, not because we have proven this approach sterilizes the joint, but because these patients have often done well after one operation [7,16], and because a stable reconstruction may let the host manage what we cannot fully clear.

The same distinction applies to the antibiotics that follow. The oral route has not been shown to be worse in the shoulder, and the intravenous route has been shown to cost the patient something [11,16].

The bottom line: we are treating the whole patient, not the presumed pathogen.

Bugs are just a part of the situation
Lewis' Woodpecker

Follow on twitter/X: https://x.com/RickMatsen


Follow on facebook: https://www.facebook.com/shoulder.arthritis

Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133//


References

1. Ahsan ZS, Somerson JS, Matsen FA 3rd. Characterizing the Propionibacterium load in revision shoulder arthroplasty: a study of 137 culture-positive cases. J Bone Joint Surg Am. 2017 Jan 18;99(2):150-154. doi:10.2106/JBJS.16.00422. PMID: 28099305.

2. Pottinger P, Butler-Wu S, Neradilek MB, Merritt A, Bertelsen A, Jette JL, Warme WJ, Matsen FA 3rd. Prognostic factors for bacterial cultures positive for Propionibacterium acnes and other organisms in a large series of revision shoulder arthroplasties performed for stiffness, pain, or loosening. J Bone Joint Surg Am. 2012 Nov 21;94(22):2075-2083. doi:10.2106/JBJS.K.00861. PMID: 23172325.

3. Bumgarner RE, Harrison D, Hsu JE. Cutibacterium acnes isolates from deep tissue specimens retrieved during revision shoulder arthroplasty: similar colony morphology does not indicate clonality. J Clin Microbiol. 2020 Jan 28;58(2):e00121-19. doi:10.1128/JCM.00121-19. PMID: 31645372.

4. Matsen FA 3rd, Butler-Wu S, Carofino BC, Jette JL, Bertelsen A, Bumgarner R. Origin of Propionibacterium in surgical wounds and evidence-based approach for culturing Propionibacterium from surgical sites. J Bone Joint Surg Am. 2013 Dec 4;95(23):e1811-7. doi:10.2106/JBJS.L.01733. PMID: 24306707.

5. Butler-Wu SM, Burns EM, Pottinger PS, Magaret AS, Rakeman JL, Matsen FA 3rd, Cookson BT. Optimization of periprosthetic culture for diagnosis of Propionibacterium acnes prosthetic joint infection. J Clin Microbiol. 2011 Jul;49(7):2490-2495. doi:10.1128/JCM.00450-11. PMID: 21543562.

6. McGoldrick E, McElvany MD, Butler-Wu S, Pottinger PS, Matsen FA 3rd. Substantial cultures of Propionibacterium can be found in apparently aseptic shoulders revised three years or more after the index arthroplasty. J Shoulder Elbow Surg. 2015 Jan;24(1):31-35. doi:10.1016/j.jse.2014.05.008. PMID: 25213827.

7. Hsu JE, Gorbaty JD, Whitney IJ, Matsen FA 3rd. Single-stage revision is effective for failed shoulder arthroplasty with positive cultures for Propionibacterium. J Bone Joint Surg Am. 2016 Dec 21;98(24):2047-2051. doi:10.2106/JBJS.16.00149. PMID: 28002367.

8. Belay ES, Danilkowicz R, Bullock G, Wall K, Garrigues GE. Single-stage versus two-stage revision for shoulder periprosthetic joint infection: a systematic review and meta-analysis. J Shoulder Elbow Surg. 2020 Dec;29(12):2476-2486. doi:10.1016/j.jse.2020.05.034. PMID: 32565412.

9. Garrigues GE, Zmistowski B, Cooper AM, Green A; ICM Shoulder Group. Proceedings from the 2018 International Consensus Meeting on Orthopedic Infections: the definition of periprosthetic shoulder infection. J Shoulder Elbow Surg. 2019 Jun;28(6S):S8-S12. doi:10.1016/j.jse.2019.04.034. PMID: 31196517.

10. Garrigues GE, Zmistowski B, Cooper AM, Green A; ICM Shoulder Group. Proceedings from the 2018 International Consensus Meeting on Orthopedic Infections: management of periprosthetic shoulder infection. J Shoulder Elbow Surg. 2019 Jun;28(6S):S67-S99. doi:10.1016/j.jse.2019.04.015. PMID: 31196516.

11. Yao JJ, Jurgensmeier K, Woodhead BM, Whitson AJ, Pottinger PS, Matsen FA 3rd, Hsu JE. The use and adverse effects of oral and intravenous antibiotic administration for suspected infection after revision shoulder arthroplasty. J Bone Joint Surg Am. 2020 Jun 3;102(11):961-970. doi:10.2106/JBJS.19.00846. PMID: 32079886.

12. Warme WJ, Hsu JE. Definition of a "true" periprosthetic shoulder infection still eludes us. J Bone Joint Surg Am. 2015 Jul 15;97(14):e56. doi:10.2106/JBJS.O.00426. PMID: 26178898.

13. Fehring TK, Otero JE, Fehring KA, Curtin BM, Springer BD, Della Valle CJ, Parvizi J, Hietpas K, Ready A, Odum SM; PJI Study Group. One-stage versus two-stage exchange arthroplasty for periprosthetic joint infection: a prospective randomized trial. J Bone Joint Surg Am. 2026 Jul 15;108(14):1070-1082. doi:10.2106/JBJS.25.00713. Epub 2026 Apr 1. PMID: 41921050.

14. Schwartz AM. One-stage versus two-stage: a potential paradigm shift for patients and surgeons? Commentary on an article by Thomas K. Fehring, MD, et al. J Bone Joint Surg Am. 2026 Jul 15;108(14):1025-1026. doi:10.2106/JBJS.26.00035.

15. Fernainy C, Randhawa AS, Frederickson M, Menendez ME. Oral antibiotic therapy for shoulder periprosthetic joint infection: current and evolving concepts. JB JS Open Access. 2026 Jan 26;11(1):e25.00244. doi:10.2106/JBJS.OA.25.00244. PMID: 41589275; PMCID: PMC12826257.

16. Yao JJ, Jurgensmeier K, Whitson AJ, Pottinger PS, Matsen FA 3rd, Hsu JE. Oral and IV antibiotic administration after single-stage revision shoulder arthroplasty: study of survivorship and patient-reported outcomes in patients without clear preoperative or intraoperative infection. J Bone Joint Surg Am. 2022 Mar 2;104(5):421-429. doi:10.2106/JBJS.21.00530. PMID: 34842573.

17. Li HK, Rombach I, Zambellas R, et al.; OVIVA Trial Collaborators. Oral versus intravenous antibiotics for bone and joint infection. N Engl J Med. 2019 Jan 31;380(5):425-436. doi:10.1056/NEJMoa1710926. PMID: 30699315.

18. Kohm K, Seneca K, Smith K, Heinemann D, Nahass RG. Successful treatment of Cutibacterium acnes prosthetic joint infection with single-stage exchange and oral antibiotics. Open Forum Infect Dis. 2023 Jul 13;10(8):ofad370. doi:10.1093/ofid/ofad370. PMID: 37539065.

19. Roger PM, Assi F, Denes E. Prosthetic joint infections: 6 weeks of oral antibiotics results in a low failure rate. J Antimicrob Chemother. 2024 Feb 1;79(2):327-333. doi:10.1093/jac/dkad382. PMID: 38113545.

20. Kelly JD 2nd, Hobgood ER. Positive culture rate in revision shoulder arthroplasty. Clin Orthop Relat Res. 2009 Sep;467(9):2343-2348. doi:10.1007/s11999-009-0875-x. PMID: 19434466.

21. Piggott DA, Higgins YM, Melia MT, Ellis B, Carroll KC, McFarland EG, Auwaerter PG. Characteristics and treatment outcomes of Propionibacterium acnes prosthetic shoulder infections in adults. Open Forum Infect Dis. 2015 Dec 9;3(1):ofv191. doi:10.1093/ofid/ofv191. PMID: 26933665.

22. Mook WR, Garrigues GE. Diagnosis and management of periprosthetic shoulder infections. J Bone Joint Surg Am. 2014 Jun 4;96(11):956-965. doi:10.2106/JBJS.M.00402. PMID: 24897745.

23. Mook WR, Klement MR, Green CL, Hazen KC, Garrigues GE. The incidence of Propionibacterium acnes in open shoulder surgery: a controlled diagnostic study. J Bone Joint Surg Am. 2015 Jun 17;97(12):957-963. doi:10.2106/JBJS.N.00784. PMID: 26085527.

24. Valenzuela MM, Odum SM, Griffin WL, Springer BD, Fehring TK, Otero JE. High-dose antibiotic cement spacers independently increase the risk of acute kidney injury in revision for periprosthetic joint infection: a prospective randomized controlled clinical trial. J Arthroplasty. 2022 Jun;37(6S):S321-S326. doi:10.1016/j.arth.2022.01.060. PMID: 35090819.

25. McFarland EG, Rojas J, Smalley J, Borade AU, Joseph J. Complications of antibiotic cement spacers used for shoulder infections. J Shoulder Elbow Surg. 2018 Nov;27(11):1996-2005. doi:10.1016/j.jse.2018.03.031. PMID: 29778591.

26. Rondon AJ, Paziuk T, Gutman MJ, Williams GR Jr, Namdari S. Spacers for life: high mortality rate associated with definitive treatment of shoulder periprosthetic infection with permanent antibiotic spacer. J Shoulder Elbow Surg. 2021 Dec;30(12):e732-e740. doi:10.1016/j.jse.2021.05.005. Epub 2021 Jun 2. PMID: 34087272.

27. Sender R, Fuchs S, Milo R. Revised estimates for the number of human and bacteria cells in the body. PLoS Biol. 2016 Aug 19;14(8):e1002533. doi:10.1371/journal.pbio.1002533. PMID: 27541692.

Wednesday, July 15, 2026

Shoulder arthritis and shoulder arthroplasty - what's new (if anything)

Two articles in the recent JSES are of note.

Advanced glenohumeral osteoarthritis: the relationship between radiographic pathoanatomy and clinical presentation asks "does the x-ray tell us how the patient is doing?" The authors studied 280 shoulders with advanced glenohumeral osteoarthritis and an intact cuff, all of which went on to arthroplasty: 147 anatomic total shoulders, 81 reverses, and 52 ream and runs [1]. Every shoulder was graded before surgery by three classifications: Samilson-Prieto, Kellgren-Lawrence, and the Walch system as modified for three-dimensional imaging [3,4,5]. The authors also measured critical shoulder angle, humeral head medialization, humeral head flattening, the length of the inferior humeral neck spur, posterior decentering, glenoid version, and glenoid inclination. Then they asked whether any of these predicted how the shoulder moved, how comfortable it was, or how the patient rated his or her health.

With two exceptions, it did not. No clinically meaningful association between Samilson-Prieto grade, Kellgren-Lawrence grade, Walch type, critical shoulder angle, medialization, version, or inclination and any patient-reported outcome or quality-of-life score. The exceptions were both on the humeral side: greater flattening of the head and a longer humeral neck spur were associated with less motion.

This replicates what we reported in 2019 in 544 shoulders [6], what Kohan and colleagues reported in 256 [7], and what Kircher and colleagues reported earlier still [13]. Four groups, four cohorts, one answer: the radiographic severity of the arthritis does not tell us what the patient is experiencing.

The second paper ---Reverse and anatomic total shoulder arthroplasty for glenohumeral osteoarthritis: a propensity-matched comparison at early and midterm follow-up ‚--- asks "does the implant choice change the result?" From a single high-volume surgeon's practice, Leinweber and colleagues matched 61 anatomic total shoulders to 61 reverses, one to one, on age, sex, body mass index, preoperative ASES, preoperative forward elevation, and Walch glenoid type [2]. Notably, they matched on the very pathoanatomy that the first paper found does not predict much. All shoulders had osteoarthritis with an intact cuff. All were seen early (about two years) and at midterm (about five years).

Both groups improved a great deal, and by the same amount. More than 96% of patients in both groups reached the minimal clinically important difference for the ASES at both time points. The anatomic patients had better external rotation at the early visit (63 degrees vs. 57 degrees) and better internal rotation as well; by five years the internal rotation advantage was gone and the external rotation difference had narrowed. Complications were 3.3% in each group.

Putting these papers side by side

Consider what each study holds constant and what it lets vary.

The second paper holds the surgeon constant and varies the implant. It finds no meaningful difference in what the patient reports.

The first paper considers the variable pathoanatomy across 280 shoulders and finds that it explains almost nothing about how the patient presents.

Neither the glenoid nor the prosthesis seems to matter. (The two measures that did survive are on the humerus, and we will come to them a bit later.) If the explanation for the differences among our patients' outcomes is not in the glenoid we spend our time classifying and not in the implant we spend our time choosing, the two possibilities left standing are the patient and the surgeon.

One incidental note: the two papers used different MCID values for the same score ‚--- 16 points for the ASES in the first [8], 10.4 points in the second [9]. Both are defensible and both are published. Whether a result is "clinically important" can depend on which threshold the authors selected.


A closer look at the first paper

Eight radiographic parameters and six classification terms were each tested against eleven outcomes: 154 comparisons, with no correction for multiplicity. At the conventional threshold, roughly eight false positives are expected by chance alone. Seventeen associations were reported as significant ‚ --- more than chance alone would produce. Of those seventeen, the authors' own MCID screen disqualified nearly every one. The findings are not false; they are true but too small to act on.

Four associations survived as "clinically relevant":

Forward elevation and humeral head flattening. A coefficient of 0.56 degrees per percentage point of humeral head flattening, against an MCID of 17.1 degrees [8], requires a change in flattening of about 30 percentage points: nearly five standard deviations, and roughly three-quarters of the entire observed range.

External rotation and head flattening. About 24 percentage points, close to four standard deviations.

Neither is a quantity that could exist in a patient.

Internal rotation and head flattening. About 10 percentage points, 1.6 standard deviations, entirely attainable.

External rotation and humeral neck spur. A difference in spur length of about 22 mm, well within the observed range of 0 to 48 mm. This result replicates Kircher [13].

Those last two are real, they are reachable, and they are both on the humeral side ‚ ---where nobody has been looking ‚ --- rather than on the glenoid, where everyone has been looking. In other words the strongest signals in 280 shoulders are from the humerus;  the glenoid---the object of twenty years of classification---apparently contributes little.

A limitation of the analysis is that every shoulder in this series went on to arthroplasty. That is a range-restricted sample: each of these patients had already crossed somebody's operative threshold, so the variance in symptoms is truncated at the low end, and truncation attenuates correlation. Some of the null result is built into the sampling. The claim is not "radiographs never relate to symptoms." The actual finding is: among patients already being considered for shoulder arthroplasty, the images doe not tell you who hurts, how much, or how far the shoulder moves.

In the subgroup analysis, the concentric (Walch A) glenoids had worse pain, worse DASH, and worse ASES than the eccentric (Walch B) glenoids. By the standard we have just applied to the rest of the paper, these differences fall below the MCID and we should not make much of them‚---so we will only state: the direction runs opposite to the expectation of most surgeons. The eccentric, retroverted, posteriorly decentered glenoid is the one that most often prompts a surgeon from an anatomic reconstruction to a reverse. In this cohort those shoulders were, if anything, the more comfortable ones. That is not what much current thinking about the B glenoid would predict, and it is worth a study designed to test it rather than a subgroup that happened to find it.



A closer look at the second paper

The surgeon performed 1,310 reverses and 546 anatomic total shoulders in the study window. Of the reverses, 133 (only 10% of the original cohort) had complete clinical and outcome follow-up at both time points; 61 entered the matched analysis. Of the anatomics, 129 (only 24% of the original cohort) qualified; 61 entered the matched analysis. So the survivorship bias and the matching dramatically reduced the studied sample size, making it less relevant to the original group of patients.

Note the inclusion rule was complete outcome scores and clinical follow-up at both visits. That rule removes the patients who were revised elsewhere, who stopped answering, who were dissatisfied and did not come back, and who died. The paper then reports a 0% revision rate for the reverse and a 1.6% acromial stress fracture rate. These are not accurate incidence estimates. A revision rate cannot be calculated in a cohort whose definition requires having completed follow-up; it requires the whole exposed group and a time-to-event analysis.

Substantial clinical benefit was reached by 95.1% of anatomic patients and 80.3% of reverse patients early. That 14.8-point gap has a P value of .027 and is discussed as a real difference. At five years the gap was 13.1 points (91.8% vs. 78.7%), with a P value of .074, and was described as no significant difference.

The two gaps are nearly the same size. What changed is which side of .05 the P value fell on, in a study that states it could not perform a power analysis. A 13-point difference in substantial clinical benefit at five years has not been shown to be absent; it has merely not been shown to be present. Those are different statements, and only the second one is supported here. The gap favors the anatomic shoulder at both time points, and it deserves a larger study rather than a rounding to "similar."

When more than 96% of both groups clear the MCID, the MCID has stopped discriminating. It tells us that both operations work, which is worth knowing and is the correct headline. It cannot tell us whether they work equally well.

Revision is not a good metric

One anatomic shoulder was revised; no reverse was. That appears in the abstract as "aTSA had more revisions." A failed anatomic shoulder has somewhere to go, --- conversion to a reverse. A failed reverse does not, and the threshold for taking a painful reverse back to the operating room is far higher, because the surgeon has less to offer. A revision count compares the availability of a salvage operation as much as it compares the durability of an implant. Counting one against zero and reporting it as a durability signal asks a single patient to carry the argument.

The same asymmetry appears in how radiographic failure was judged. A reverse with baseplate lucencies and broken screws was recorded as a non-failure because the patient was comfortable. Asymptomatic glenoid lucencies after anatomic arthroplasty, graded by the Lazarus system [10], were simultaneously presented as the durability liability. One standard should apply to both.

The whole argument for using a reverse in a cuff-intact arthritic shoulder is long-term durability: glenoid loosening and late cuff failure at ten, fifteen, twenty years. Five years samples precisely the window in which the anatomic shoulder is known to do well [11], and the one series that has followed eccentric-wear shoulders past that window to a minimum of seven years found the anatomic reconstruction holding up [12]. This study cannot address the question that prompted it.

The variable neither study measured

Put the two preoperative cohorts next to each other. Both are advanced osteoarthritis with an intact cuff, in the same country, in the same era. The shoulders moved almost identically before surgery: mean forward elevation 96.6 degrees in the first series, 97.5 degrees and 98.2 degrees in the second.

The patients, however, were not in the same condition. Mean preoperative ASES was 29.9 in the first series and 41.4 and 38.9 in the second. Mean pain was 7.4 out of 10 versus 5.3 and 5.5. The patients in one practice arrived at the operating room roughly ten ASES points and two pain points better off than the patients in the other.

There is more than one way to get such a gap. These are different practices with different referral streams, different payer mixes, and different geography; the first cohort also includes 52 ream-and-runs, a self-selected group that has no counterpart in the second practice. So the difference may reflect who walks through the door as much as when the surgeon decides to operate. But that distinction does not rescue the number. Referral pattern and operative threshold are both properties of the practice, not of the shoulder. Either way, what varies between these two cohorts is the surgeon's context, not the patient's pathoanatomy ‚ --- and it varies by more than the pathoanatomy does. This is the kind of unwanted variation that Kahneman called noise.

What this may mean

If you are a patient looking at your own x-ray or CT scan, the first paper is reassuring: the severity of what you see on the film does not predict how much your shoulder will hurt, how far it will move, or how you will feel about your life. Some very ugly-looking shoulders belong to comfortable people, and some mild-looking ones hurt badly. The x-ray describes the joint. It does not describe you.

If you are a patient choosing between an anatomic and a reverse replacement for arthritis with an intact cuff, the second paper says that at five years, in the hands of one experienced surgeon, both do well. The anatomic shoulders showed somewhat better rotation early, though by the standard applied above that early difference is at the edge of what a patient would notice. What happens after five years is not yet known, and that is the question that actually separates these two operations.

If you are a surgeon, the two papers together take away two of the things you were counting on. The images you spend the most time studying did not explain how the patient presented. The implant you spent the most time deciding on did not explain how the patient ended up. What is left is the judgment that sits between them: whom you offer an operation to, when in the course of the disease you offer it, what you tell the patient to expect, and how well you execute. Those are the variables neither study measured, and the ten-point preoperative ASES gap between these two practices is a direct measurement of how much they vary from one surgeon to the next.

We keep looking for the explanation in the imaging and in the implant because those are the things we can see and buy. The evidence keeps pointing somewhere else.

Both of these are careful papers by serious groups, and both are more honest in their limitations sections than most. The first calls itself a pilot and names its multiplicity problem. The second calls itself hypothesis-generating, names the risk of type II error, and says that longer follow-up is needed. 

The larger point stands. 

Twenty years of classifying the glenoid has not produced a classification that predicts what the patient feels. 

Ten years of moving the reverse into the cuff-intact arthritic shoulder has not yet produced evidence that it does better than the anatomic shoulder we already had. 

The trial that could settle it would randomize, follow to fifteen years, and report substantial clinical benefit rather than the MCID.

Which leaves the question: if the picture does not explain the patient's presentation and the implant does not explain the result, what does?



Where are we going?
Snow Geese
Skagit County




[1] Covarrubias O, Luther L, Portnoff B, Levins J, Hoffman R, Molla V, Toavs T, Molino J, Paxton ES, Green A. Advanced glenohumeral osteoarthritis: the relationship between radiographic pathoanatomy and clinical presentation. J Shoulder Elbow Surg 2026;35:1620-1631. https://doi.org/10.1016/j.jse.2026.01.007

[2] Leinweber KA, Bowler AR, Diestel DR, McDonald-Stahl M, Arnold RP, Le K, Dunn WR, Kirsch JM, Jawa A. Reverse and anatomic total shoulder arthroplasty for glenohumeral osteoarthritis: a propensity-matched comparison at early and midterm follow-up. J Shoulder Elbow Surg 2026;35:1632-1641. https://doi.org/10.1016/j.jse.2025.12.008

[3] Samilson RL, Prieto V. Dislocation arthropathy of the shoulder. J Bone Joint Surg Am 1983;65:456-460.

[4] Kellgren JH, Lawrence JS. Radiological assessment of osteo-arthrosis. Ann Rheum Dis 1957;16:494-502.

[5] Bercik MJ, Kruse K 2nd, Yalizis M, Gauci MO, Chaoui J, Walch G. A modification to the Walch classification of the glenoid in primary glenohumeral osteoarthritis using three-dimensional imaging. J Shoulder Elbow Surg 2016;25:1601-1606. https://doi.org/10.1016/j.jse.2016.03.010

[6] Matsen FA 3rd, Whitson A, Hsu JE, Stankovic NK, Neradilek MB, Somerson JS. Prearthroplasty glenohumeral pathoanatomy and its relationship to patient's sex, age, diagnosis, and self-assessed shoulder comfort and function. J Shoulder Elbow Surg 2019;28:2290-2300. https://doi.org/10.1016/j.jse.2019.04.043

[7] Kohan EM, Hill JR, Lamplot JD, Aleem AW, Keener JD, Chamberlain AM. Severity of glenohumeral osteoarthritis does not correlate with patient-reported outcomes. J Shoulder Elb Arthroplast 2020;4:2471549220901873. https://doi.org/10.1177/2471549220901873

[8] Simovitch RW, Elwell J, Colasanti CA, Hao KA, Friedman RJ, Flurin P, et al. Stratification of the minimal clinically important difference, substantial clinical benefit, and patient acceptable symptomatic state after total shoulder arthroplasty by implant type, preoperative diagnosis, and sex. J Shoulder Elbow Surg 2024;33:e492-e506. https://doi.org/10.1016/j.jse.2024.01.040

[9] Levy JC, Everding NG, Gil CC Jr, Stephens S, Giveans MR. Speed of recovery after shoulder arthroplasty: a comparison of reverse and anatomic total shoulder arthroplasty. J Shoulder Elbow Surg 2014;23:1872-1881. https://doi.org/10.1016/j.jse.2014.04.014

[10] Lazarus MD, Jensen KL, Southworth C, Matsen FA 3rd. The radiographic evaluation of keeled and pegged glenoid component insertion. J Bone Joint Surg Am 2002;84:1174-1182. https://doi.org/10.2106/00004623-200207000-00013

[11] Kirsch JM, Puzzitiello RN, Swanson D, Le K, Hart PA, Churchill R, et al. Outcomes after anatomic and reverse shoulder arthroplasty for the treatment of glenohumeral osteoarthritis: a propensity score-matched analysis. J Bone Joint Surg Am 2022;104:1362-1369. https://doi.org/10.2106/JBJS.21.00982

[12] Cuff DJ, Simon P, Patel JS, Munassi SD. Anatomic shoulder arthroplasty with high side reaming versus reverse shoulder arthroplasty for eccentric glenoid wear patterns with an intact rotator cuff: comparing early versus midterm outcomes with minimum 7 years of follow-up. J Shoulder Elbow Surg 2023;32:972-979. https://doi.org/10.1016/j.jse.2022.10.017

[13] Kircher J, Morhard M, Magosch P, Ebinger N, Lichtenberg S, Habermeyer P. How much are radiological parameters related to clinical symptoms and function in osteoarthritis of the shoulder? Int Orthop 2010;34:677-681. https://doi.org/10.1007/s00264-009-0846-6
 

Tuesday, July 7, 2026

Two year outcomes for reverse total shoulder for osteoarthritis and the effect of survivorship bias.

How survivorship bias can affect reports of two-year outcomes

When we quote a two-year success rate for an operation, the figure typically describes the outcomes for the patients who came back to be assessed at two or more years after surgery.

Those who did not return are not a random sample of all the patients having the surgery. Patients lost to follow-up tend to have fared worse than those who complete follow-up, because the reasons they drop out are often themselves adverse — death, revision, or disappointment with an early result that leads them to transfer their care elsewhere [1, 2, 3].

Excluding patients who have done poorly prior to the two-year mark — considering only those who remain for the two-year analysis — makes the operation look better than it actually was, a phenomenon known as survivorship bias [4].

A model of survivorship bias

Here is an illustrative hypothetical example of a thousand patients having an RSA for cuff-intact osteoarthritis with two-year follow-up.

In the first six months patients were lost as follows: 20 shoulders revised, 5 patients dead, 15 transferred to another practice, and 40 who simply stopped responding — 80 in all. Of note, certain problems occur early in the postoperative period: acute infection, instability, acromial fracture, and life-limiting frailty.

During the second six months there were 7 more revisions and 8 deaths, 25 transfers, and 110 who did not respond for unknown reasons — 150 in the interval.

The second-year losses are almost entirely loss of contact rather than loss of the shoulder: 3 revised and 12 more dead, 35 gone to other practices, and 200 more who stopped answering — 250 in that interval.

By the time of the two-year analysis, 30 of the 1,000 have been revised, 25 have died, and 75 have transferred their care; 350 more have gone quiet without a recorded reason. That is 480 lost from view, none of whom are included in the two-year analysis of results; only 520 remain available for study.

Figure 1. A hypothetical cohort of 1,000 patients followed over two years. At each interval, patients drop out of view for four kinds of reasons, and the mix changes as time passes. Of the 520 analyzed at two years, 9 out of 10 report success; but out of all 1,000 operated on, a successful outcome is documented for only about 4.7 out of every 10.

The four reasons carry different weight, and their proportions change over the two years. Revisions come early: RSA’s dominant early failures — acute infection, instability, and acromial fracture — appear mostly in the first months, so revisions decline from 20 in the first half-year to 3 in the second year. Deaths run the other way, accumulating with time in an elderly group, from 5 to 8 to 12 across the intervals. Transfers of care, and above all patients who simply stop responding, grow steadily as contact is lost — from 40 silent patients in the first six months to 200 in the second year. By the time of the two-year analysis, loss of contact, not loss of the shoulder, accounts for most of the missing.


Applying the model of survivorship bias to actual data on two-year outcomes for RSA for osteoarthritis

The model becomes meaningful when anchored to what the literature reports. Two published figures set the scale.

The first is the rate of two-year follow-up. In a multicenter shoulder arthroplasty registry, only about half of patients — about 5 out of 10 — provided two-year patient-reported outcomes [5]; registries that send repeated reminders might improve the return to  8 out of 10.

The second is the success rate among those who do return: in high-volume single-surgeon series of RSA for this diagnosis, about 9 out of 10 report being better and satisfied [6, 7].

Consider the combined effect of these two rates. Of the 520 with a known two-year result, 9 out of 10 — 468 patients — report success; but across the full 1,000 who had the surgery, those 468 provide an overall documented success rate of only about 4.7 out of 10.

Had the follow-up rate reached 8 out of 10 rather than 5, about 800 would have a known result, and the same 9 out of 10 would give about 720 documented successes — roughly 7 out of 10.

The span from 4.7 to 7 out of 10 is set entirely by the follow-up rate.

The 4.7 out of 10 is the floor. It counts every patient without a documented success as a non-success, so that the true whole-cohort success rate (if it could be known) might be higher than 4.7 out of 10. 

The missing patients are not a random subset of the cohort. Those lost to revision, death, or a transfer of care each had reason to fare worse than the returners. The larger unresponsive group has, on average, done less well than those who answer — though that association is weaker and less certain than for the other types of losses [1, 2, 3].

Reporting the returners’ 9 out of 10 as the rate for the overall cohort makes the operation look better than the data support [4].

However,  a successful outcome is only documented for about 5 of every 10  patients considering all those having the surgery, not 9 out of 10.

A fuller accounting would mean following the patients who leave — the revised, the transferred, and above all the ones who quietly stop answering — well enough to know how they actually did. Until a series does that, the appropriate two-year estimate for successful outcome for RSA in osteoarthritis is somewhere above the documented 4.7 out of 10 rate for the entire cohort and below the rate of 9 out of 10 considering only the patients returning for two year analysis. 

And two years is the favorable case. The effect of survivorship bias grows as follow-up lengthens: over five, ten, and fifteen years, follow-up falls further and the number of missing patients grows, so the distance between the returners’ success rate and the whole-cohort success rate only widens. The longer the follow-up a reported success rate claims, the greater the effect of survivorship bias.

It's about survivorship


Bald eagles
Montlake Cut


Follow on twitter/X: https://x.com/RickMatsen
Follow on facebook: https://www.facebook.com/shoulder.arthritis

References

[1] Solberg TK, Sørlie A, Sjaavik K, Nygaard ØP, Ingebrigtsen T. Would loss to follow-up bias the outcome evaluation of patients operated for degenerative disorders of the lumbar spine? A study of responding and non-responding cohort participants from a clinical spine surgery registry. Acta Orthop. 2011;82(1):56–63.

[2] Murnaghan ML, Buckley RE. Lost but not forgotten: patients lost to follow-up in a trauma database. Can J Surg. 2002;45(3):191–195.

[3] Torrens C, Martínez R, Santana F. Patients lost to follow-up in shoulder arthroplasty: descriptive characteristics and reasons. Clin Orthop Surg. 2022;14(1):112–118.

[4] Elston DM. Survivorship bias. J Am Acad Dermatol. Published online June 18, 2021.

[5] Patel M, Sekar MG, McDaniel L, Kisana HM, Sykes JB, Amini MH. Changes from baseline in patient-reported outcomes and patient satisfaction do not vary significantly between 1 and 2 years postoperatively after shoulder arthroplasty: a multicenter analysis of 2580 patients. Semin Arthroplasty JSES. 2025;35(2):235–245.

[6] Puzzitiello RN, Moverman MA, Glass EA, Swanson DP, Bowler AR, Le K, Kirsch JM, Lohre R, Jawa A. Clinically significant outcome thresholds and rates of achievement by shoulder arthroplasty type and preoperative diagnosis. J Shoulder Elbow Surg. 2024;33(7):1448–1456.

[7] Ahmed AF, Glass EA, Swanson DP, et al. Predictors of poor and excellent outcomes following reverse shoulder arthroplasty for glenohumeral osteoarthritis with an intact rotator cuff. J Shoulder Elbow Surg. 2024;33(6S):S55–S63.


Thursday, July 2, 2026

Robotics in Shoulder Arthroplasty: ASES Podcast 156.

Robotics in Shoulder Arthroplasty is a timely and informative ASES podcast hosted by Peter Chalmers and Brian Waterman with two guests: Michael Freehill and Vahid Entezari.

Here is a summary of the points the participants made. Where the guests referred to a specific study, a link to the primary source has been added so that readers can consult it directly and form their own view.

The episode presented enthusiasm tempered by candor about the current state of evidence. Both guests described themselves as early adopters, and both acknowledged that outcome data specific to robotic shoulder arthroplasty is not yet available. Their case for the technology rested primarily on precision in executing a preoperative plan, the ability to generate intraoperative data, enhanced capability in complex and deformed anatomy, and anticipated — as yet undemonstrated — long-term benefit. They also included a discussion of the market and workflow factors that shape adoption.

Here are some details.

Dr. Entezari suggested that shoulder arthroplasty has two distinct steps — planning and execution — and that while preoperative planning has advanced considerably, precision execution has historically lacked dedicated tools. He referenced previous work on the accuracy of planning and patient-specific instrumentation (PSI) showing that preoperative planning and PSI tightened the spread of deviation from the preoperative plan in comparison to manual techniques. See Accuracy of 3-Dimensional Planning, Implant Templating, and Patient-Specific Instrumentation in Anatomic Total Shoulder Arthroplasty. 

From this he drew two conclusions: that surgeons doing manual-only surgery may underestimate their own deviation (“you don’t know what you don’t know”), and that enabling technologies — navigation, PSI, or robotics — become especially relevant in cases with significant deformity or difficult exposure. These technologies have the potential to reduce substantial deviation from the preoperative plan.

Dr. Freehill built on this from a training perspective, recalling high-volume leaders he trained under questioning their own accuracy: whether their humeral head cuts and glenoid version were as accurate as they desired. He described preoperative planning as a major advance and PSI as a further step for deformed anatomy, while noting that guides still have limitations related to soft tissue and exposure. He also emphasized that most shoulder arthroplasties are performed by lower-volume surgeons — those doing fewer than roughly ten a year — and suggested these surgeons might benefit from tools that help with execution.

Both guests placed the clearest value of robotics in complex cases: significant glenoid deformity (such as B2 and B3 morphology and higher degrees of retroversion), revision situations, and cases requiring bone graft or augment preparation.

Dr. Entezari described a reverse arthroplasty for arthropathy that had developed after a dislocation and fracture, with anterior glenoid bone loss, in which the robot was used to shape a humeral head autograft to fit the glenoid face, offering it as an example of capabilities that could extend to graft and augment work.

The guests pointed to potential new capabilities: a robotically controlled burr that can sculpt bone at angles a saw or reamer cannot achieve, the potential for subscapularis-sparing and minimally invasive approaches. They pointed out that robotics could capture potentially useful data such as how precisely the preoperative plan was carried out.

The hosts pointed to the mixed hip and knee literature, noting meta-analyses that have not shown a long-term benefit in clinical outcomes even though the rate of use of robotics was rapidly increasing.

The guests did not claim superiority in patient reported outcomes for robotics in shoulder arthroplasty. Dr. Entezari suggested that malpositioning may not affect these outcomes but that adherence to the preoperative plan may matter for long-term implant survival.

Dr. Freehill cited NICE’s assessment of robot-assisted orthopaedic surgery. The relevant NICE document is its early value assessment of robot-assisted surgery for hip and knee replacement (HealthTech guidance HTG743in which patient-reported outcomes and complications were similar between robotic and conventional surgery and implant alignment was consistently more precise with robotics. The revision data were limited, and the committee was uncertain whether more precise alignment yields better clinical outcomes. He pointed to the need for head-to-head studies, while acknowledging such studies are difficult and slow to complete.

Dr. Entezari cited a Monte Carlo simulation whose authors concluded that, when value is judged by tail-risk rather than by mean accuracy, the principal contribution of robotic assistance in shoulder arthroplasty is risk containment — with clinical and economic implications that remain unproven to date — and that further research linking its use to patient outcomes is warranted and may ultimately redefine the technology’s value proposition.  See Mako Robotic-Assisted Glenoid Preparation in Reverse Shoulder Arthroplasty: A Tail-Risk Reduction Perspective Compared with Manual and Patient-Specific Guide Techniques.

With respect to the learning curve for robotics, Dr. Freehill cited experience suggesting an inflection point around ten cases, with time efficiency approaching break-even, and drew an analogy to hip and knee arthroplasty, where operative time tends to normalize with experience. 

[Parenthetically, we can observe that the inflection point of ten cases is close to the annual case volume of a “low-volume surgeon.”]

Both guests suggested the real barrier to adoption is neither time nor cost but workflow change. 

Dr. Freehill discussed how the technology is presented to hospital administrators, whom he characterized as focused on time-zero costs rather than on anticipated ten- to fifteen-year outcomes or on accuracy data. He noted that an administrator is unlikely to be swayed by work comparing robotic glenoid preparation favorably with patient-specific instrumentation and, still more, with manual technique, because their attention stays on the bottom line. He acknowledged the numbers work against robotics: citing the existing literature, he said that breaking even by minimizing revisions would require avoiding roughly one revision in fifteen (about 7%), whereas revision rates in reverse arthroplasty are currently under 2%, so on that basis one is, in his words, already arguing against the cost-benefit case. See  Is Robotic-Assisted Reverse Shoulder Arthroplasty Economically Justified? A Break-Even Analysis.  He suggested the more persuasive pitch to an administrator combines several elements: that a robot from certain companies is often already in the hospital for hip and knee procedures; that adopting cutting-edge technology is anticipated to improve market share; and that greater accuracy can be framed as serving the patient’s long-term interest.

Dr. Entezari added that hospitals in a region compete for the same patients, and that device companies invest heavily in direct-to-consumer advertising, so patient demand — already seen in hip and knee — becomes part of the rationale. He noted that costs should fall as more surgeons use the technology. He reiterated that outcome data are not yet available and that early adopters have the opportunity to generate the needed data.

The guests contrasted navigation and robotics. Navigation has a smaller footprint, lower cost, and a lower barrier to entry. Robotics offers greater precision, but with a larger and more expensive early-generation factor, and with current systems limited in the procedures they can perform and the range of implants they can accommodate. 

Dr. Entezari suggested neither overselling nor underselling the technology — telling patients that it makes the surgery more precise while being candid that its effect on their individual outcome is not known. Dr. Freehill described a mix of curiosity about cutting-edge technology and hesitation about a non-human element in the operation, and said he uses clinic time to explain the surgeon’s central role.

Both guests urged caution about how the technology is learned. They argued that younger surgeons should first become proficient in manual arthroplasty, comparing the situation to skills that have faded elsewhere in orthopaedics, such as freehand pedicle screw placement, manual knee cuts, and open shoulder stabilization. 

Dr. Entezari described losing registration during a case after contact with a tracker and having to rely on his manual skills to complete the operation. Dr. Freehill echoed the concern that trainees might become proficient with the technology without ever learning the underlying skills of standard shoulder arthroplasty.

Conclusion: this was a balanced presentation that emphasized the potential role of robotics in transferring a preoperative plan to the patient in the operating room, without claiming that robotic shoulder arthroplasty has been associated with improved patient outcomes.


Which is better?

Tundra Swans
Lake Washington