A colleague wrote recently about a recent break-even analysis of robotic-assisted reverse total shoulder arthroplasty (RSA) with a straightforward question: "where did the modeled range of $500 to $3,000 per case in a recent article 1 come from? Were those list prices, or institutional costs?"
This seemingly straightforward question opened the proverbial Pandora’s Box.
So, hang on for a bit of a wild ride.
1. What problem are we trying to solve?
Before asking what a technology costs, it is worth asking what it is for. These two very different questions are routinely run together, and they lead to different places.
The first is: how can the cost of this technology be justified? That question starts with the technology and looks for a rationale for its use. It is the question a break-even analysis answers, and it is the one most of the literature is now busy with.
The second is: what measures would improve the clinical outcomes of my patients, by how much, and what would they cost to implement in my hospital? That question starts with the patients and looks for a remedy. Only the second can tell a surgeon whether to buy anything, because only the second begins with a problem rather than with a product.
There is a way to answer the second question that costs nothing and can be done this week. A surgeon looks back over her own experience with reverse total shoulder arthroplasty and lists the outcomes that disappointed her or her patients — the shoulders that stayed painful, the ones that never regained the lost function the patient came for, the instability, the glenoid loosening, the infections, the reoperations. Then she asks of each one, case by case, whether an enabling technology designed to give a component placement that more accurately reproduces a preoperative plan would plausibly have prevented it.
If the answer is not a high percentage of them, the technology may be a non-starter at any cost. If the answer is a substantial share, then the cost question becomes worth the considerable trouble that follows. Either way the audit comes first, and it requires no vendor, no capital committee, and no observer with a stopwatch.
This is Codman’s End Result Idea applied to a purchasing decision: know what happened to your own patients before deciding what to buy. Everything that follows in this piece concerns what happens after such an audit comes back positive.
2. Why value, and not cost?
Cost alone has never settled anything in medicine. Consider a technology that restored the function of a failing kidney. It could be enormously expensive and we would still judge it worth having, and we would not spend much time arguing about the per-case disposables. What we would want to know is how much function it restored, in how many patients, and for how long. Only against that answer does the cost become interesting.
Value is a ratio. The numerator is what the patient gains; the denominator is what it costs to deliver. Our discussion of shoulder technology has concentrated almost entirely on the denominator, and a denominator without a numerator cannot produce a judgment about worth.
Both halves have the same shape, and each can be stated in one sentence.
Term | Definition |
Numerator | Lifetime comfort and function for this patient with the intervention, minus lifetime comfort and function for the same patient without it |
Denominator | Preoperative plus intraoperative plus postoperative cost with the intervention, minus preoperative plus intraoperative plus postoperative cost without it |
Divide the second into the first and the result is dollars per unit of patient benefit — which is, in plain language, the incremental cost-effectiveness ratio that health economists have been using for decades. The definitions are simple in concept. Every difficulty in this post comes from trying to obtain the four quantities they require.
Some features of those definitions deserve particular attention.
The word “lifetime” is critical: the numerator is not the score at two years but the accumulated comfort and function over the patient’s remaining life, which is why durability matters more than the early result. The three-phase structure of the denominator maps exactly onto a time-driven activity-based costing process map, which is why that method is the right instrument for it. And the second term of each difference — the same patient without the intervention —cannot be known. For any given patient we get one arm and only one. Both halves therefore require randomization, matching, or modeling, and neither escapes that requirement. Notice also that both terms are written about a particular patient. The definitions are patient-specific by construction, which means a single number describing the generalized value of a technology will never exist. Furthermore, the average value of a technology in one medical center (say a large specialty urban hospital) will not be the value of the same technology in another (say a smaller hospital in a rural setting).
The definitions also show why certain events count twice. A pin-site fracture adds to the postoperative cost and subtracts from lifetime comfort and function; it makes value worse in both terms. An averted episode of instability does the reverse, adding to the numerator and subtracting from the denominator. Events of this kind move the ratio about twice as far as their cost alone would suggest, in whichever direction they fall.
It would be convenient to say that cost is the half we can count and benefit is the half we cannot. That is not true. Both halves are hard, and they are hard in different ways. The numerator is hard because it takes years: comfort and function have to be collected before surgery and at fixed intervals afterward, and durability cannot be rushed. The denominator is hard for reasons patience does not fix. The money is set in confidential agreements between hospital and vendor and cannot be discovered. Operating room time is measurable, but planning time is not being measured at all. Added people and added space are shared and lumpy, and do not divide by case. A CT scan counts only against a baseline practice that plans using only plain films, which means the analysis has to say which baseline it used. And while some of the technology’s complications can be counted directly, the ones it shares with conventional surgery require a controlled comparison.
One category refuses to stay on either side. Where the two approaches share a complication — infection, instability, loosening — the technology may cause it or prevent it: same measurement, unknown sign, and a study that measures one half without the other cannot place it. The complications unique to the technology are different, and simpler. They can only be a cost.
The asymmetry in our current attention does damage in both directions. Weighed against an unmeasured benefit, any cost looks unjustifiable. Presented without its full costs, any technology looks cheap. Both errors are avoidable, and neither is fixed by counting the easy things.
The asymmetry also shapes what the existing analyses can conclude. At an incremental cost of $1,500 per case and a revision treatment cost of $22,920, robotic RSA reaches economic neutrality only by preventing one revision for every 15 procedures — an absolute risk reduction of 6.55%.1 Set that against the revision rate actually reported: 2.1% within the first postoperative year in United States hospital billing data, or about one revision in every 48 patients.2 To be effective the technology would have to prevent roughly three times as many revisions as actually are occurring. Break-even is unreachable on that endpoint, not because the required reduction is large in absolute terms, but because it exceeds the entire event rate available to be reduced. The finding is not that the technology is too expensive. It is a finding that the benefit being counted is too small a thing to count. If the only good the technology is credited with is preventing a 2.1% event, almost nothing can be shown to be worth its cost.
None of this is abstract. Enabling technologies are entering shoulder practice now.3 Capital committees are approving purchases, service lines are being built, and patients are being consented. Those decisions are being made with or without careful data, and if surgeons do not supply the analysis it will be supplied by whoever has a number at hand.
3. How do we measure the benefit?
This is the half that is currently missing altogether. Five things could serve as the numerator, and they are not equivalent.
What we could count | What it tells us | What it leaves out |
Component accuracy | Whether the preoperative plan was executed as intended | Whether the patient is any better when the technology is used. |
Revision avoided | A hard, costly event prevented | Rare — about 2.1% in the first year — so the total benefit available to be captured is small |
Time to revision | Durability; delay has value even when the endpoint is eventually reached | Requires long follow-up, which no shoulder robotic cohort yet has |
Comfort and function at the MCID | What the patient came to the office for | Requires prospective collection before surgery and at fixed intervals afterward |
Predictability and fewer outliers | A possible benefit to the surgeon and the system | Not yet linked to any patient-reported gain |
Component position has been compared between robotic, manual, and patient-specific approaches in the laboratory and in systematic review.4,5 These studies demonstrate accuracy in carrying out a preoperative plan, which is a statement about execution rather than about the patient. Accuracy is a reasonable thing to want, but it is not by itself a benefit to the patient.
Revision is a hard endpoint, cheap to count, and a surrogate. Its rarity is precisely what makes it a poor numerator: there is not enough of it to justify much cost, however effective the technology might be. In addition we know that absence of a revision does not indicate that the patient's outcome is satisfactory (see All the Failures We Cannot See)
Comfort and function are what the patient came for, and this is the numerator that could justify a large expenditure if the technology delivers it. The thresholds already exist. In a multicenter cohort of 5,851 arthroplasties, minimal clinically important difference values were 2.1 points for the Simple Shoulder Test and 13.9 for the ASES score, with substantial clinical benefit thresholds of 3.8 and 33.1.6,7 An analysis claiming effectiveness should report the proportion of patients reaching those thresholds and how long the gain lasts with and without the application of the enabling technology.
Collecting this is not conceptually difficult. It requires asking every patient, before surgery and at fixed intervals afterward, how the shoulder is doing, and reporting the answers whichever way they fall. That is the Codman End Result Idea, and it is a century old. It is also the only route to a numerator large enough to make an expensive technology worth buying.
4. How do we measure the incremental cost?
This half is not simple either. Some of it has a proven method and is simply not being applied; some can be measured only against information a hospital is not permitted to see; and some cannot be established at a single institution at all.
Start with the right definition
The cleanest definition of the cost of the technology is the total cost of caring for this patient with it minus the total cost of caring for the same patient without it. Everything that differs counts; everything that does not, drops out.
A second distinction matters. An implant is a single line: the cost of the new component minus the cost of what it replaces; if the same implant is used with and without the technology, this term cancels out. On the other hand, an enabling technology is a bundle of different components, each of which adds to the overall cost. The cost of robotics is not only the cost of the robot.
Personnel and time
This part has a proven method, but it is costly to carry out. Time-driven activity-based costing builds a process map of every preoperative, intraoperative, and postoperative activity, and a single observer follows consecutive patients from admission to discharge with a chronometer, recording minutes for each professional role. Applied to 25 consecutive same-day primary RSAs performed with conventional instrumentation, it produced an average standardized hospital cost of $16,547, against $21,579 for the same patients by traditional costing; a minute of labor across all personnel categories cost about $24.8
Measuring the increment means running the same map on the technology arm at the same institution and subtracting. Two refinements are needed. First, the map must be extended beyond the hospital day to capture preoperative planning, which happens in the surgeon’s office, generates no charge, and rarely appears in published costing studies of shoulder technology. Second, operative time must be reported in blocks rather than as a single mean, because it changes with experience: in 30 consecutive Mako robotic RSAs by one fellowship-trained surgeon, mean operative time fell from 105.5 minutes in the first ten cases to 67.8 in the third ten, crossing the center’s conventional benchmark of 74.9 minutes at about case 23.9 A single mean across an adoption series describes neither the learning period nor the steady state.
Added personnel and space confound per-case measurement altogether. A trained technologist, a console footprint, storage, and reprocessing capacity are lumpy and shared. They are better handled as fixed costs divided by annual volume, alongside capital — which will be covered below.
Induced care
Anything the patient receives only because the technology requires it belongs in the increment. Preoperative CT is the obvious case, with its charge, its scheduling burden, and its radiation. A structured review of technology-assisted knee arthroplasty found that studies reporting robotic or patient-specific instrumentation costs frequently omitted imaging altogether, and added a fixed CT charge of $420 where it was missing.10 Whether that increment applies in the shoulder depends on baseline practice: in the TDABC series above, conventional cases used no patient-specific guides and no other enabling technology, but a CT-based preoperative plan was available in the operating room.8 If CT is obtained anyway for glenoid morphology, the incremental imaging cost of adding a robot is small. If the preoperative planning is carried out using plain films rather than a CT, the technology has created a scan that would not otherwise exist. Either answer is defensible; leaving it unstated is not.
Complications, in two classes
Complications can be divided into two types.
The first kind are unique to the technology. Conventional RSA has no tracker pins, no robotic arm, and no possibility of converting to a manual technique partway through the procedure. The conventional rate of these events is therefore zero by construction, which has a useful consequence: a pin-site fracture, a pin-site infection, a tracker dislodgement, an intraoperative hardware failure, an unexpected arm movement during preparation, or a case abandoned mid-procedure is attributable to the technology without any comparison arm at all. The excess equals the gross incidence. The sign is not in doubt. And these can be counted prospectively in a single-arm series from the first case onward, which makes them the most tractable item in the entire denominator.
What such events look like in practice is already on record. In robotic-assisted knee arthroplasty, among 263 adverse event reports over an eighteen-month period, 204 involved total knee arthroplasty, the most frequent being unexpected robotic arm movement during bone cuts (59, 28.9%), inaccurate bone cuts (26, 12.7%), and fluid leak or residue on robotic components (25, 12.3%).11 In robotic-assisted hip arthroplasty, 521 reports describing 546 discrete events showed intraoperative hardware failure as the commonest problem (304, 55.7%) and inaccurate cup placement second (63, 11.5%); the robot was abandoned in 13.0% of reports, and a surgical delay occurred in 28%, averaging 17.9 minutes and running as long as an hour.12 Across 15,004 navigated and robotic knee arthroplasties, pin-related complications occurred in 0.95%, with pin-site fractures in 0.16% presenting on average 8.5 weeks after surgery and requiring treatment from protected weight-bearing to plate fixation or intramedullary nailing.13 An earlier review put the overall pin complication rate at 1.4%.14
Costing these is arithmetically easy once the events are known. An abandoned robot is a conversion. A 17.9-minute delay is operating room time at roughly $24 a minute in personnel alone.8 A pin-site fracture treated surgically is a reoperation costing far more than any disposable it offsets.
Two cautions belong with those figures. Passive reporting databases have no denominator, so they describe what happens rather than how often; the pin review is the exception, with 15,004 cases behind its 0.95%. And all of the evidence comes from the hip and knee. Shoulder robotics typically use a glenoid tracking array and carries the same possibility of conversion, so the category exists in the shoulder by construction — but the incidence has not been reported.
The second kind are the complications the two approaches share: infection, instability, glenoid loosening, revision for any cause. Here the conventional rate is not zero, so the gross incidence tells us nothing and a controlled comparison is required. Here too the sign is genuinely unknown — if better positioning reduces instability, the term is negative and belongs in the numerator rather than the denominator. This is the part that cannot be settled without a comparison arm, and it is worth separating from the part that can.
Infection deserves separate mention because added operative time raises the question directly. Does the technology increase operative time? If so, is the increment the same for all surgeons or for an individual surgeon across her or his learning curve? Operative time is associated with infection after shoulder arthroplasty, with an inflection point reported at 180 minutes15 and a relative risk of 1.24 for any complication per additional 20 minutes,16 and a meta-analysis of 427,361 primary knee arthroplasties found higher infection odds beyond 90 and 120 minutes.17 But longer cases are also harder cases, so duration may be a marker rather than a mechanism — and in the shoulder the dominant organism enters from transected dermal sebaceous glands at the moment of incision, despite skin preparation and intravenous prophylaxis, an inoculum delivered at the start of the case rather than accumulated over its length.18
What can and cannot be captured
The categories are not equally tractable, and the difficulty is not random: the items easiest to measure are also the ones most often reported.
What differs | How it can be captured | Where it stands |
Operating room time | Direct observation of consecutive cases against the same center’s conventional benchmark, reported in blocks so the learning curve is visible | Measured in one single-surgeon shoulder series. Likely to vary surgeon to surgeon. |
Personnel time | Extend the TDABC process map to the technology arm and re-time every role, then cost at the same per-minute salary rates | Method established; not yet applied to a technology comparison in the shoulder. Costly to measure. May change with volume and experience. |
Planning time | Add preoperative planning to the process map; it happens outside the operating room and generates no charge | Not reported in any costing study of shoulder technology |
Added personnel and space | Treat as fixed cost and divide by annual volume, as with capital — a technologist, console footprint, storage, reprocessing capacity | Absorbed into overhead at one center, a capital project at another |
Induced imaging | Count only the scans that would not have been obtained anyway, which requires stating baseline practice | Usually unstated |
Complications unique to the technology | Count them directly — the conventional rate is zero, so no comparison arm is needed; cost them as conversions, delay minutes, and reoperations | Reported in hip and knee series; not yet counted in any shoulder cohort |
Complications shared with conventional surgery | The excess over conventional, which needs a controlled comparison | No shoulder data; sign of the net effect genuinely unknown |
Costs under a bundled agreement | The difference between what the hospital now pays and what it would have paid otherwise | Governed by confidential terms; generally not knowable |
Three of these are measurable and simply are not being measured. Three cannot be settled at a single institution, and one cannot be reported at all. A point estimate that omits the difficult categories has not been conservative; it has set them to zero.
There is a way to turn that limitation into an argument rather than a weakness. Instead of pursuing a complete total, ask what benefit is already required by the costs a hospital can document — capital amortization and disposables alone. Every omitted category can only raise that bar, so a conclusion drawn at the floor is robust to precisely the uncertainty that is hardest to resolve. A defensible floor is worth more than a false total.
5. What method puts the two halves together?
For the denominator: TDABC, not traditional costing
Traditional costing averages overhead across patients and overstates the result. The two methods applied to the same 25 shoulder arthroplasties differed by about $5,000, with personnel at $4,483 by traditional costing against $1,483 by TDABC.8 In hip and knee arthroplasty the discrepancy has approached a factor of two.19,20 Prior shoulder work using TDABC established that implants and personnel dominate, with modest variation between hospitals and surgeons.21,22A technology compared against a traditionally costed baseline is being compared against an inflated number. The caveat is that TDABC costs can be expected to vary with volume and experience within a given center; furthermore these data cannot be applied to other centers – the are no generalizable.
Two limits on the method deserve stating, because they rarely are. The first is that TDABC is itself expensive. It requires an observer following patients through an entire hospital day, analysts to cost the activities, and institutional support to do any of it. A hospital that cannot afford the technology generally cannot afford the study that would tell it whether to buy the technology, which means the evidence gets generated where it is least needed.
The second is that the answer is not stable. Salaries, staffing ratios, care pathways, and vendor agreements change from year to year, so a benchmark is a snapshot with a short shelf life. It also depends on who operates. A benchmark generated when attending surgeons commit to performing the entire procedure themselves will not describe the same institution’s usual practice, in which residents and fellows perform a substantial share of the case. That difference changes both operative time and personnel cost, and it changes them most in the arm where the learning curve is steepest.
For the numerator: not a trial, at least not first
A simulation-based power analysis in the knee estimated that detecting revision-rate reductions attributable to robotic or navigated technology would require very large numbers of patients.23 That is the honest reason two alternatives exist.
Break-even analysis, introduced in the shoulder to evaluate vancomycin powder,24 asks how far an event rate must fall before an added cost is repaid. Its output is a target, not a finding: at $1,500 per case it required an absolute risk reduction of 6.55%, and the highest incremental cost compatible with break-even was about $481.1 Simulation asks which practice settings could clear a cost-effectiveness threshold;25 in the knee, median reductions in lifetime revision risk ranged from 0.9% to 5.1%, and robotic assistance cleared the threshold only in larger practices within narrow bands of utilization.10
Neither demonstrates that a technology delivers a benefit. Both convert an argument about enthusiasm into an arithmetic statement of the benefit that would have to exist, and both are available before any trial reports. A paper using them should say which inputs were measured and which were assumed. The knee simulation, to its credit, used the most favorable published precision values and said outright that its estimates may have been overly optimistic.10
Of the available frameworks, the quality-adjusted life year is the one that can actually express the intuition behind the kidney example, because it values health gained rather than events avoided. The knee simulation used it.10 No published analysis of shoulder robotics has, and none can until comfort and function are collected prospectively.
Test the assumptions that actually move the answer
Sensitivity analysis is often run on the wrong parameter. In the robotic RSA break-even model, varying the baseline revision rate did not change the required risk reduction at all; it stayed at 6.55%. What moved the result was the cost of the failure being prevented, falling from 15% at a $10,000 revision to 3% at $50,000 and 1.5% at $100,000.1 Report the extrema as well as the median. A reporting standard for all of this already exists: the CHEERS 2022 statement sets out a 28-item checklist for health economic evaluations.26 Orthopaedics has been slow to adopt it.
A workable sequence
· Decide the numerator before starting, and say whether it is a surrogate.
· Build the conventional TDABC benchmark at your own institution.
· Run the same process map on the technology arm at the same institution, in blocks, and subtract.
· Add induced imaging, planning time, and fixed costs divided by your actual annual volume.
· Put your own contracted numbers into the equation and see what benefit your hospital would require.
· Collect comfort and function on every patient in both arms, before surgery and at fixed intervals afterward.
· Replicate across centers of different size and different contracts.
6. How can we estimate what we can never observe?
Both definitions require a quantity that does not exist: the same patient, treated the other way. For any individual we get one arm and only one, and no amount of careful measurement recovers the other. This is not a peculiarity of surgical technology. It is the central problem of clinical research, and it has standard partial solutions.
The move that makes it tractable is to stop asking for the value in a patient and ask instead for the value in a circumstance — a defined stratum of patients, at a stated institution, under a stated contract. Individual treatment effects are not identifiable by any method. Average effects within a stratum are estimable, under assumptions that can be written down and argued with. That is a smaller claim than we might want, but it is the one available.
Approach | What it can give | What it costs or assumes |
Randomized trial | The clean answer, with no assumption about assignment | Impractical for revision as the endpoint; far smaller if the endpoint is comfort and function |
Within-patient comparison | Both arms in the same person — the closest thing to observing the counterfactual | Requires bilateral disease and staged surgery; small numbers; sides are not perfectly comparable |
Concurrent matched cohort | A stratum-level estimate at one institution, with both arms under the same cost structure | Confounding by indication is the central threat; the assignment rule must be documented, not assumed |
Before-and-after at one center | Controls the institution, the surgeons, and the contract | Learning curve and secular trend are mixed into the treatment effect |
Quasi-random assignment | Assignment driven by robot or room availability rather than by the patient | Only credible if availability is genuinely unrelated to patient characteristics |
Modeled counterfactual | An answer now, from a measured conventional benchmark plus modeled effects | Every assumption must be stated; this is what break-even and simulation already do |
Threshold reasoning | What the benefit would have to be, without estimating what it is | Needs no counterfactual at all, and answers a narrower question |
The first row deserves more attention than it usually gets. The argument that a trial is impractical rests on the revision endpoint: detecting a difference in an event occurring in roughly 2% of patients requires very large numbers.23 But that is a statement about the endpoint, not about the design. A difference in the proportion of patients reaching the minimal clinically important difference in comfort and function is a far commoner event, and a trial powered on it is correspondingly smaller. If we chose the numerator we actually care about, the study we keep calling infeasible might not be.
Where a trial is genuinely out of reach, the discipline that makes observational work answer a causal question is to specify the trial one would have run — eligibility, assignment, endpoint, follow-up — and then emulate it in a registry, rather than analyzing whatever data happen to exist. The main threat is confounding by indication, and in this setting it runs in both directions: a surgeon may reach for the technology in difficult anatomy, which biases against it, or reserve it for straightforward cases during adoption, which biases for it. The assignment rule has to be documented rather than assumed away, and matching should be on the variables that move both terms of the ratio — glenoid morphology and bone loss, age, sex, comorbidity, and baseline comfort and function.
Like for like
Whichever approach is chosen, the whole enterprise rests on one requirement: the two arms have to be alike in everything except the intervention. That sounds obvious and is where most of the available evidence falls down.
What must match | Specifically | What goes wrong if it does not |
The patient | Diagnosis, glenoid morphology and bone loss, cuff status, age, sex, comorbidity, and baseline comfort and function | Room to gain differs between the arms; unmatched severity biases the result in whichever direction the selection ran |
The surgeon | The same surgeon, or the same level of experience, and both arms past the learning curve — or the curve reported separately | Adoption-period time and complications get attributed to the technology rather than to inexperience |
The implant | The same prosthesis in both arms | The comparison becomes robot-plus-implant-A against manual-plus-implant-B, which answers a different question |
The pathway | The same anesthesia and block protocol, subscapularis management, and discharge model | Postoperative cost and early function differ for reasons that have nothing to do with the technology |
The institution and period | The same building, contracts, salaries, and overhead, with overlapping dates | Cost differences reflect the contract or the calendar year rather than the intervention |
The measurement | The same costing method in both arms, the same outcome instruments at the same intervals, and the same definitions of complication and revision | A measured cost set against a modeled one is not a comparison |
Judged against that standard, none of the work discussed above qualifies, and each falls short on a different item — which is not a criticism, since none of them claimed to. The TDABC benchmark has no technology arm at all; it was built to be the conventional half of a comparison that has not yet been made.8 The operative-time series compares robotic cases against a conventional benchmark at the same center, which matches institution and surgeon but not patients.9Break-even sets a measured cost against a modeled benefit.1 Simulation matches nothing by design; it projects.10 The gap in the literature is not a missing conclusion. It is a missing like-for-like comparison.
This requirement is also the argument for doing the work inside a single institution first. A prospective registry running both arms concurrently holds the institution, the contracts, the pathway, and the measurement constant automatically; only the patients then need matching, which is the one item on that list that statistical methods can address. There is a cost to that convenience, and it should be stated openly: the more perfectly matched the comparison, the narrower the setting to which it applies. A flawless single-center result answers the question for that center. Both things are true, and neither is a reason to skip the other.
Two further points make this more workable than it first appears. First, when the same question is asked at several centers, the between-center differences should be reported as a result rather than averaged away as noise; a value that varies from $1,800 to $4,300 per case across institutions is itself the finding, and it is the finding a hospital needs in order to locate itself. Second, one piece of the denominator escapes the problem entirely. The complications unique to the technology have a conventional rate of zero, so they need no counterfactual and can be counted in a single arm beginning with the first case.
None of this yields the value for a given patient. It yields the value for a stratum, at a center, with assumptions stated plainly enough to be disputed. That is what clinical epidemiology has always delivered, and it is enough to decide with — provided we say which stratum, which center, and which assumptions.
7. Why the answer differs — between hospitals, and between patients
What volume does, and does not, change
Part of an enabling technology’s cost is fixed and part of it is not, and the distinction matters more than it is usually given credit for. Disposables, imaging, added operating room time, and the people in the room recur with every case and do not fall as volume rises. Capital, service contracts, software subscriptions, training, and dedicated space are spread across whatever caseload there is. Only the second group behaves the way the arithmetic below assumes, and how large that group is depends entirely on the agreement. Where a vendor supplies the platform in exchange for a component commitment, there is little fixed cost left to amortize at all. Where the vendor does not, a large capital outlay has to be spread across more cases, and the pressure to do so pushes the technology toward straightforward cases that do not need it — so the arithmetic of amortization becomes an argument for using the technology on the patients least likely to benefit from it. The figures below isolate the capital component alone, holding the recurring per-case costs constant, to show what volume does to that part. Using the mean values from the structured review of technology-assisted knee arthroplasty — annual capital of $139,000 and per-case variable costs of $1,490 for robotic systems10 — and holding the revision treatment cost at $22,920:1
Cases per year | Capital per case | Plus per-case variable | Required risk reduction | Number needed to treat |
50 | $2,780 | $4,270 | 18.6% | 5 |
100 | $1,390 | $2,880 | 12.6% | 8 |
200 | $695 | $2,185 | 9.5% | 11 |
400 | $348 | $1,838 | 8.0% | 13 |
600 | $232 | $1,722 | 7.5% | 13 |
The same technology costs about two and a half times as much per case at 50 cases a year as at 400. It is also worth noticing where the widely used $1,500 base case sits: it corresponds almost exactly to the per-case variable component with the capital burden excluded. Adding realistic capital at plausible volumes puts the true increment between roughly $1,700 and $4,300 — in the upper half of the modeled $500 to $3,000 range or above it. That does not weaken the published conclusion; it strengthens it, since a higher increment makes break-even harder to reach.
The contract layer on top
Volume also determines the cost a hospital is charged, and here the effect compounds rather than cancels. The incremental costs at the center of every one of these models are set by confidential agreements between vendor and hospital. A high-volume center may receive a platform and its service at no charge in return for a component commitment — the Gillette model, give away the razor and sell the blades. A lower-volume hospital cannot negotiate that deal.
Non-disclosure compounds it further. A hospital cannot learn what a comparable hospital pays for the identical implant, so it cannot judge whether its own cost is competitive; a reader of a cost analysis cannot tell whether the figure used resembles anything their institution would face; and the vendor is the only party to the negotiation who knows the full distribution of costs.
Consider the common arrangement in its fullest form: platform and technologist provided at no charge in exchange for a commitment of x cases using the vendor’s implant, with the implant discounted as well. Run the counterfactual definition mechanically and the added cost comes out negative — the technology appears to pay the hospital to use it. That is arithmetically correct and economically false. The capital cost has not disappeared; it has been converted into a volume commitment and shifted into the implant charge, across a bundle that is not disclosed. The discount is purchased with forfeited negotiating leverage and foreclosed implant choice. A volume commitment creates a standing institutional reason to use the technology, which can push it toward the marginal patient rather than the difficult anatomy where it might help most. And a center operating under such an agreement has a financial interest in the technology’s success that belongs in any disclosure, yet cannot be disclosed because the terms are confidential. A negative increment is not good news; it is a signal that the accounting boundary has been drawn in the wrong place.
Because the required benefit is simply the incremental cost divided by the value of the outcome improved, it scales directly with whatever a hospital pays:1
Incremental cost per case | Required absolute risk reduction | Number needed to treat | Contracting scenario |
$0 | 0% | Any benefit qualifies | Supplied under a volume agreement |
$481 | 2.1% | 48 | Published break-even ceiling |
$1,500 | 6.5% | 15 | Published base case |
$3,000 | 13.1% | 8 | Upper end of the modeled range |
$7,000 | 30.5% | 3 | No volume offset |
Two surgeons could read identical clinical evidence, apply an identical model, and reach opposite and equally correct conclusions. Value is not a property of a device. It is a property of three things together: a device, a contract at a particular institution, and a particular patient. Change any one and the answer changes.
The learning curve is heavier at low volume too
If operative time returns to benchmark at about case 23,9 a practice doing 400 cases a year absorbs the adoption period in about three weeks. A practice doing 50 absorbs it over half a year, during which a larger share of its total caseload carries the penalty. The same is true of added personnel and space: absorbed into existing overhead at a large center, a capital project at a small one.
And it differs from patient to patient
Institutional variation is only half of the generalizability problem. Both terms of the ratio are written about a particular patient, and both vary substantially across the patients in any practice.
The numerator varies with how much there is to gain. A severely deformed or eroded glenoid offers more room for precision to matter than a concentric one, and a younger patient has a longer remaining life over which any durability advantage accumulates — the lifetime integral is simply larger. The denominator varies too: time-driven activity-based costing in the shoulder has shown higher costs in patients with greater comorbidity and worse preoperative pain and function.21 The same technology is therefore both more valuable and more expensive in different patients, and not always the same ones.
The knee simulation quantified this, and the effect is large. In a lower-risk population, the median reduction in lifetime revision risk attributable to robotic assistance was 1.8%, with cumulative gains of about 1.1 quality-adjusted life years per 100 patients. In an elevated-risk population — differing only in age, body mass index, and sex — the same technology produced a median reduction of 4.6% and about 4.0 QALYs per 100 patients. The cost-effectiveness consequences followed: robotic assistance cleared the threshold in the lower-risk population only in practices of at least 600 cases a year and within a narrow band of utilization, but in the elevated-risk population it could clear the same threshold in practices as small as 100 cases a year.10 Nothing about the technology changed. Only the patients did.
That argues for selective use rather than blanket adoption, and it raises two difficulties worth stating. The first is ethical: the authors of that simulation noted that prioritizing patients by age and sex, which is what their own risk model implies, could reasonably be seen as discriminatory.10 The second is structural. A vendor agreement requiring a committed number of cases pushes an institution toward using the technology on everyone, at precisely the moment the analysis argues for using it on a selected few. The contract and the evidence pull in opposite directions.
The benchmark does not generalize either
The same limitation applies to the baseline. The TDABC benchmark discussed above was produced at a single high-volume institution by four fellowship-trained surgeons who committed to performing the procedures themselves rather than guiding trainees, with same-day discharge; the authors state plainly that their results may not extrapolate to other surgeons.8 That is a carefully measured number, and it is the cost of that operation in that building.
This is a property of single-institution designs, not a criticism of the institutions doing the work. The center that goes first necessarily reports its own conditions, and a high-volume referral practice is the right place to begin. But the hospital in Omak does not have four high-volume shoulder surgeons, will not be charged the same implant cost, and will not reproduce those personnel timings. The remedy is not to fault the single-center study; it is to extend the method across centers of different size and different contracts, and to use a standardized cost framework so that the figures can be compared at all.27
There is an uncomfortable inversion at the end of this. The surgeons who might gain most from standardized execution are those performing the fewest of these operations, and they practice in exactly the settings where the technology costs the most and break-even is furthest out of reach.
What a reader should ask — and who could possibly do it
For those building one of these analyses, and for those reading one:
The question | What the analysis should make explicit |
What is the benefit being claimed? | The numerator named in advance, with a plain statement of whether it is a surrogate |
How much benefit counts as benefit? | A threshold anchored to what patients perceive, not to statistical significance or to degrees of version |
Whose dollars? | Perspective — hospital, payer, patient, or society — stated and held consistent |
Compared with what? | The same patient treated without the technology, at the same institution, costed by the same method |
Which object is being costed? | The implant increment and the enabling-technology increment separated and reported on their own |
What is inside the increment? | Platform amortization and service, per-case disposables, induced imaging, incremental room time, planning time, added personnel and space, and device-specific complications — with any category set to zero said so explicitly |
At what volume? | The annual case number over which fixed costs were divided, and the amortization period assumed |
What is the endpoint? | Named in advance, with a plain statement of whether it is a surrogate |
Why this method? | The reason a trial was not feasible, including the sample size one would require |
Which assumptions move the result? | Sensitivity analysis over the parameters that actually change the answer, with extrema alongside medians |
To whom does it apply? | Case volume, surgeon experience, patient risk profile, utilization rate, and whether a volume commitment exists |
Who benefits from the answer? | Full conflict disclosure and the source of every cost figure used |
A fair objection to that list: who has the energy? Almost no one. It is a standard for judging a published claim, not homework for a practicing surgeon, and nobody outside a research unit is going to work through twelve items before a purchasing meeting.
The short version for the surgeon boils down to three questions.
What problem am I trying to solve? What share of my own disappointing outcomes would this technology plausibly have prevented? And what does it cost in my hospital, under my hospital’s agreement? The first two cost nothing and are the ones that decide most cases. The third requires only that someone ask the contracting office for a number.
There is a larger difficulty behind the objection. A full value analysis can be performed, at substantial expense, at a large institution, and it may well show that the technology is worth its cost there. That result describes that institution. It does not establish the value of the technology across the reverse shoulder arthroplasties performed each year in the United States and elsewhere, most of them outside high-volume referral centers, under different agreements, with different caseloads and different surgeons. An answer obtained where the study was affordable describes the setting where the study was affordable. Which is why the surgeon’s own audit is not a poor substitute for the formal analysis — for most practices, it is the only version of the question that can actually be answered.
Where this leaves shoulder robotics today
Demonstrated: robotic and other digital technologies improve the accuracy of glenoid preparation and positioning.4,5 The conventional operation costs roughly $16,500 by TDABC at a high-volume center.8 The operative time penalty resolves after about two dozen cases in one single-surgeon series.9
Not demonstrated: that the technology improves comfort or function by a minimal clinically important difference; that it reduces revision at any time point; that accuracy gains persist at ten years; what any given hospital actually pays; or what the technology’s own complications cost in the shoulder. The numerator, in other words, is unmeasured.
None of that is evidence that the technology does not help. Absence of evidence is not evidence of absence. It is a description of work that has not yet been done, much of it difficult for structural reasons rather than for want of effort — which is why it deserves to be done carefully, and why the results will take years rather than months.
A question to close on
So the question is not whether shoulder robotics is expensive. It is the one we started with: what problem are we trying to solve, and would this solve it? Answer that first, from your own results. Then, and only then, is it worth finding out what it would cost you.
How can we cut the best deal for our patients?
Follow on facebook: https://www.facebook.com/shoulder.arthritis
Follow on LinkedIn: https://www.linkedin.com/in/rick-matsen-88b1a8133/
References
1. Menendez ME, Moverman MA, Schiffman CJ, Matsen FA 3rd.
Is robotic-assisted reverse shoulder arthroplasty economically justified? A break-even analysis. J Shoulder Elbow Surg. 2026;35(9):2313-2317. doi:10.1016/j.jse.2026.03.010
2. Corso KA, Smith CE, Vanderkarr MF, Debnath R, Goldstein LJ, Varughese B, et al. Postoperative revision, complication and economic outcomes of patients with reverse or anatomic total shoulder arthroplasty at one year: a retrospective, United States hospital billing database analysis. J Shoulder Elbow Surg. 2025;34:e59-e71. doi:10.1016/j.jse.2024.05.009
3. Sanchez-Sotelo J. Robot-assisted shoulder arthroplasty. JSES Int. 2025;9(3):974-980. doi:10.1016/j.jseint.2025.02.004
4. Athwal GS, Nelson A, Antuna S, Ponce B, Mighell M, St Pierre P, et al. Glenoid preparation in reverse shoulder arthroplasty: robotic arm-assisted preparation compared to manual preparation and patient-specific guides. J Shoulder Elbow Surg. 2025;34:2022-2030. doi:10.1016/j.jse.2024.12.007
5. Lee D, Yoo J, Yoon JP, Oh KS, Chung SW. Comparison of patient-specific instrumentation, navigation, and mixed reality technologies for accurate glenoid positioning in reverse total shoulder arthroplasty: a systematic review and meta-analysis. J Shoulder Elbow Surg. 2025. doi:10.1016/j.jse.2025.07.019
6. Simovitch RW, Elwell J, Colasanti CA, Hao KA, Friedman RJ, Flurin PH, Wright TW, Schoch BS, Roche CP, Zuckerman JD. Stratification of the minimal clinically important difference, substantial clinical benefit, and patient acceptable symptomatic state after total shoulder arthroplasty by implant type, preoperative diagnosis, and sex. J Shoulder Elbow Surg. 2024;33(9):e492-e506. doi:10.1016/j.jse.2024.01.040
7. Simovitch R, Flurin PH, Wright T, Zuckerman JD, Roche CP. Quantifying success after total shoulder arthroplasty: the substantial clinical benefit. J Shoulder Elbow Surg. 2018;27(5):903-911. doi:10.1016/j.jse.2017.12.014
8. Barret H, Vazquez A, Dholakia R, Borah BJ, Barlow JD, Morrey ME, Sperling JW, Sanchez-Sotelo J. The hospital cost of primary reverse total shoulder arthroplasty performed with traditional instrumentation: a time-driven activity-based costing (TDABC) benchmark analysis in the era of digital enabling technology. J Shoulder Elbow Arthroplast. 2026 (in press). doi:10.1016/j.jsea.2026.100087
9. Buac NP, Chhokar M, Menendez ME. Robotic-assisted reverse shoulder arthroplasty achieves operative time neutrality after an initial learning period. Int Orthop. 2026;50(4):853-857. doi:10.1007/s00264-026-06774-7
10. Hickey MD, Masri BA, Hodgson AJ. Can technology assistance be cost effective in TKA? A simulation-based analysis of a risk-prioritized, practice-specific framework. Clin Orthop Relat Res. 2023;481(1):157-173. doi:10.1097/CORR.0000000000002375
11. Pagani NR, Menendez ME, Moverman MA, Puzzitiello RN, Gordon MR. Adverse events associated with robotic-assisted joint arthroplasty: an analysis of the US Food and Drug Administration MAUDE database. J Arthroplasty. 2022;37(8):1526-1533. doi:10.1016/j.arth.2022.03.060
12. Graefe SB, Kirchner GJ, Pahapill NK, Nam HH, Dunleavy ML, Haines N. Adverse events associated with robotic-assistance in total hip arthroplasty: an analysis based on the FDA MAUDE database. Hip Int. 2024;34(6):688-694. doi:10.1177/11207000241263315
13. Di Carlo G, Zampogna B, Criseo N, Aragona D, Pugliesi O, Calaciura S, Fenga D, Sanzarello I, Leonetti D. The dark side of precision: pin-related complications in computer-navigated and robotic-assisted knee arthroplasty. J Clin Med. 2026;15(10):3793. doi:10.3390/jcm15103793
14. Thomas TL, Goh GS, Nguyen MK, Lonner JH. Pin-related complications in computer navigated and robotic-assisted knee arthroplasty: a systematic review. J Arthroplasty. 2022;37(12):2291-2307.e2. doi:10.1016/j.arth.2022.05.012
15. Schmitt MW, Chenault PK, Samuel LT, Apel PJ, Bravo CJ, Tuttle JR. The effect of operative time on surgical-site infection following total shoulder arthroplasty. J Shoulder Elbow Surg. 2023;32(11):2371-2375. doi:10.1016/j.jse.2023.05.010
16. Wilson JM, Holzgrefe RE, Staley CA, Karas S, Gottschalk MB, Wagner ER. The effect of operative time on early postoperative complications in total shoulder arthroplasty: an analysis of the ACS-NSQIP database. Shoulder Elbow. 2021;13(1):79-88. doi:10.1177/1758573219876573
17. Shin KH, Kim JH, Han SB. Greater risk of periprosthetic joint infection associated with prolonged operative time in primary total knee arthroplasty: meta-analysis of 427,361 patients. J Clin Med. 2024;13(11):3046. doi:10.3390/jcm13113046
18. Matsen FA 3rd, Butler-Wu S, Carofino BC, Jette JL, Bertelsen A, Bumgarner R. Origin of Propionibacterium in surgical wounds and evidence-based approach for culturing Propionibacterium from surgical sites. J Bone Joint Surg Am. 2013;95(23):e181. doi:10.2106/JBJS.L.01733
19. Akhavan S, Ward L, Bozic KJ. Time-driven activity-based costing more accurately reflects costs in arthroplasty surgery. Clin Orthop Relat Res. 2016;474(1):8-15. doi:10.1007/s11999-015-4214-0
20. Palsis JA, Brehmer TS, Pellegrini VD, Drew JM, Sachs BL. The cost of joint replacement: comparing two approaches to evaluating costs of total hip and knee arthroplasty. J Bone Joint Surg Am. 2018;100(4):326-333. doi:10.2106/JBJS.17.00161
21. Menendez ME, Lawler SM, Shaker J, Bassoff NW, Warner JJP, Jawa A. Time-driven activity-based costing to identify patients incurring high inpatient cost for total shoulder arthroplasty. J Bone Joint Surg Am. 2018;100(23):2050-2056. doi:10.2106/JBJS.18.00281
22. Carducci MP, Gasbarro G, Menendez ME, Mahendraraj KA, Mattingly DA, Talmo C, et al. Variation in the cost of care for different types of joint arthroplasty. J Bone Joint Surg Am. 2020;102(5):404-409. doi:10.2106/JBJS.19.00164
23. Hickey MD, Anglin C, Masri B, Hodgson AJ. How large a study is needed to detect TKA revision rate reductions attributable to robotic or navigated technologies? A simulation-based power analysis. Clin Orthop Relat Res. 2021;479:2350-2361.
24. Hatch MD, Daniels SD, Glerum KM, Higgins LD. The cost effectiveness of vancomycin for preventing infections after shoulder arthroplasty: a break-even analysis. J Shoulder Elbow Surg. 2017;26(3):472-477. doi:10.1016/j.jse.2016.07.071
25. Dubois RW. Cost-effectiveness thresholds in the USA: are they coming? Are they already here? J Comp Eff Res. 2016;5(1):9-11.
26. Husereau D, Drummond M, Augustovski F, de Bekker-Grob E, Briggs AH, Carswell C, et al. Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) statement: updated reporting guidance for health economic evaluations. BMJ. 2022;376:e067975. doi:10.1136/bmj-2021-067975. (Also Value Health. 2022;25(1):3-9. doi:10.1016/j.jval.2021.11.1351)
27. Visscher SL, Naessens JM, Yawn BP, Reinalda MS, Anderson SS, Borah BJ. Developing a standardized healthcare cost data warehouse. BMC Health Serv Res. 2017;17(1):396. doi:10.1186/s12913-017-2327-8
Disclosure: The author has no financial relationships with any manufacturer of orthopaedic devices, and is a co-author of one of the studies cited.
