AlphaSteel › Blog › What The Placebo Arms Did
Trial readingWhat the Placebo Arms Did: Reading Vitality and Prostate Trials From the Control Group Out
In the large saw palmetto trials, the men on placebo got better. In the biggest, the urinary symptom score fell by 2.99 points on placebo and by 2.20 on saw palmetto over 72 weeks, and about 44 per cent of the placebo group had a fall of three points or more. Tongkat ali trials measure mood scales, sexual questionnaires and hormone levels instead of urinary scores, and their control arms show the same habit: people given a look-alike capsule change too.
That is why a before-and-after gain on a supplement, taken alone, proves little. The only figure worth reading is the gap between the two arms, and the control arm tells you how big a gain has to be before it counts as one.
- Placebo arms improved. In the 72-week CAMUS trial the placebo group’s symptom score fell 2.99 points against 2.20 for saw palmetto, and 44.2 per cent of the placebo group had a fall of three points or more, against 42.6 per cent on the extract.
- It is common across prostate trials. A meta-analysis of 25 trials found the symptom score improving by 4.4 points on average at 12 months in placebo or sham arms, while urine flow did not.
- Selection alone moves scores. Healthy men retested 10 to 20 days apart, after applying typical trial entry rules, improved without any treatment.
- Read your own result with care. Take more than one baseline, use a standard scale, change one thing at a time, and treat a gain of a few points as inside what placebo arms show.
Most write-ups of a supplement trial begin with the treatment arm: what the extract did, by how much, in whom. Try starting at the other end. The men who were given a look-alike capsule with nothing active inside were measured too, on the same scales, over the same weeks, and what happened to them says a great deal about what a change in the treatment arm can mean.
This post takes that view of two bodies of research that sit behind the AlphaSteel label: the saw palmetto trials in men with urinary symptoms and the tongkat ali trials in men with vitality and sexual well-being outcomes. It is not a review of whether saw palmetto or tongkat ali works, which the row-by-row posts cover, and it does not draw on customer reviews or anyone’s before-and-after story about a product. It is about what a control group is for, and what a reader can and cannot take from a before-and-after.
Why start with the control group
A trial has two jobs. One is to see whether people on the treatment change. The other, which matters more, is to see whether they change more than they would have anyway. The control arm exists to answer the second job, and everything that is not the treatment lives in it: the natural ups and downs of the condition, the men’s expectations, the attention of the study staff, the tendency of unusually high scores to drift back towards average, and any other change in their lives over the same weeks.
Two terms are worth keeping apart. The placebo response is whatever happens in the placebo arm, for any reason. The placebo effect is the part caused by receiving the placebo itself. A large placebo response does not mean the placebo effect is large. Much of it may be regression to the mean and ordinary fluctuation, which would have happened with no capsule at all. The 2010 Cochrane review of placebo interventions, discussed below, makes that distinction the centre of its method.
For a reader of a supplement label the practical point is short. A supplement page that says men improved by so much after taking the capsule is quoting a treatment-arm number. The number that matters is the gap between the arms, and to know how big that gap has to be, you need to see what the control arm did.
What the placebo arms did in the big saw palmetto trials
Two large randomised trials, both funded by the US National Institutes of Health (as the NCCIH page on saw palmetto notes) and both double-blind, tested saw palmetto against placebo in men with urinary symptoms. The existing saw palmetto post reports what each concluded. Here the interest is in the control groups.
The Saw Palmetto Treatment for Enlarged Prostates (STEP) trial randomised 225 men over 49 with moderate-to-severe symptoms to 160 mg of saw palmetto extract twice a day or placebo for one year (Bent 2006). The abstract reports the between-group figure and not each arm’s own change: no significant difference in the American Urological Association Symptom Index (a mean difference of 0.04 points, 95 per cent confidence interval −0.93 to 1.01) and none in peak urine flow, prostate size, residual urine, quality of life or PSA. A gap of 0.04 points means whatever the placebo arm gained, the extract arm gained the same.
The Complementary and Alternative Medicine for Urological Symptoms (CAMUS) trial, run later, randomised 369 men aged 45 and over, with a symptom index between 8 and 24, to saw palmetto at one, then two, then three times the standard dose, or placebo, over 72 weeks (Barry 2011). The numbers by arm are these.
| Baseline | At 72 weeks | Change | |
|---|---|---|---|
| Saw palmetto (up to 960 mg a day) | 14.42 | 12.22 | −2.20 |
| Placebo | 14.69 | 11.70 | −2.99 |
| Difference in change | 0.79 points in favour of placebo |
Scores are the AUA Symptom Index, range 0 to 35, higher is worse. From the abstract of Barry 2011.
Look at the placebo row first. Men who took nothing active saw their scores fall by about three points over 72 weeks. The trial’s own planning shows why that matters: its sample-size calculation treated a two-point difference as roughly what men with baseline scores of 8 to 19 report as “slight” improvement. On that yardstick the placebo arm improved by more than a slight amount. The proportion of men whose score fell by three points or more was 42.6 per cent on saw palmetto and 44.2 per cent on placebo, and at 72 weeks the men’s own global ratings averaged 3.6 on the extract and 3.5 on placebo, between “a little better” and “about the same”.
A before-and-after reading of the saw palmetto arm alone would say the extract worked, since scores fell by 2.2 points. The comparison says the same fall happened, slightly larger, with placebo. The Cochrane review of 27 trials, which the formula-level evidence post covers, reached the same conclusion for saw palmetto alone (Franco 2023).
Was the blind kept, and who dropped out
A control group only works if the men do not know which arm they are in. If they can tell, expectations differ between the arms, and the comparison is contaminated in one direction or the other. CAMUS did something many supplement trials skip: it tested the blind. Saw palmetto extract has a mild odour, so the capsules were blister-packed to avoid unblinding during pill counts, and at the end of the study men were asked to guess their assignment (Barry 2011).
| Arm (men who answered) | Thought saw palmetto | Thought placebo | Not sure |
|---|---|---|---|
| Saw palmetto (149) | 45 (30.0%) | 67 (45.0%) | 37 (24.8%) |
| Placebo (154) | 39 (25.3%) | 66 (42.9%) | 49 (31.8%) |
The two sets of answers did not differ significantly (P = 0.36). From the full text of Barry 2011.
Men on the extract were, if anything, more likely to guess they were on placebo than on saw palmetto. So the blind held, which is part of why the trial’s result is hard to dismiss.
Dropout is the other quiet lever. CAMUS randomised 369 men, analysed 357 in its main analysis, and 306 completed the full 72 weeks. For men who left early, the researchers estimated the missing scores by statistical imputation, and they confirmed the result in a per-protocol analysis of the 306 who stayed. Both analyses pointed the same way. That is a good sign, because what happens to the men who leave, and how the analysis handles them, can tilt a small trial without anyone intending it to.
The Lopatkin trial of a saw palmetto and nettle root product (see the formula post) shows another design feature: after a two-week single-blind placebo run-in, 257 men were randomised, and over 24 weeks the symptom score fell by 6 points on the product and by 4 on placebo (Lopatkin 2005). Even in a trial that favoured the product, the placebo arm improved by two thirds as much as the treated arm.
Placebo response across prostate trials: what 25 trials found
One large trial can be unusual, so it helps to see the pattern across many. Eredics and colleagues pooled randomised trials in men with urinary symptoms that had a placebo or sham arm and 12 months of follow-up (Eredics 2017). They found 25 trials with 10,587 men: 23 with a placebo arm (four of plant extracts, nine of 5-alpha-reductase inhibitors, five of alpha-blockers, three of combination drug therapy and two of an intraprostatic injection) and two with a sham procedure.
| Type of trial | Mean change, IPSS points |
|---|---|
| All 25 trials | −4.4 (range 0.7 to 6.8) |
| Plant extracts | −3.6 |
| 5-alpha-reductase inhibitors | −3.4 |
| Alpha-blockers | −4.3 |
| Combination drug therapy | −4.3 |
| Intraprostatic injection | −3.9 |
| Sham microwave heat treatment | −6.8 |
IPSS is the International Prostate Symptom Score, range 0 to 35, lower is better. From the abstract of Eredics 2017.
Two things stand out. The first is size: on average a placebo or sham arm improved by more than four points at a year, and in the plant-extract trials by 3.6. The second is the split between what men reported and what was measured. Peak urine flow in the same arms rose by only 0.8 mL/s overall, which the authors call not relevant, and in the plant-extract trials it moved by −0.3 mL/s, meaning no improvement. The score, which is the man’s own report of his symptoms, changed a good deal, while the instrument reading did not.
The authors add that the size of the placebo effect varied considerably between studies even at 12 months. So there is no single “placebo number” to subtract. Each trial has to be read against its own control arm, and a before-and-after with no control arm cannot be corrected by borrowing someone else’s.
A single-arm example shows how this plays out. In the trial that compared the saw palmetto and nettle root product with finasteride, the symptom score fell from 11.3 to 8.2 at 24 weeks in the herbal group, a fall of about three points (Sökeland and Albrecht 1997). That trial had no placebo arm. Set beside the placebo arms above, a three-point fall is squarely within the range that men on placebo show, which is why the trial could say the two active treatments were equivalent but could not say either one was better than nothing.
Regression to the mean, the change that needs no belief
Some of every placebo arm’s improvement needs no explanation about expectation at all. Regression to the mean is a statistical phenomenon in which unusually high or low measurements tend to be followed by measurements closer to the average. Barnett and colleagues put it plainly: it can make natural variation in repeated data look like real change, it is more noticeable when measurement is noisy, and it is more noticeable still when follow-up is examined only in a sub-sample selected on a baseline value (Barnett 2005).
Trials of men with urinary symptoms are built the way that produces it. They admit men whose scores are above a threshold, so they select men on a high reading, and a man is likely to have been having a bad stretch when he qualified. Sech and colleagues tested this directly. They gave 145 men with no known prostate disease, mean age 52, the AUA Symptom Index and a flow recording twice, 10 to 20 days apart, with no treatment at all (Sech 1998). Scores on the two occasions were well correlated (0.73 to 0.89), but not identical. When the researchers applied typical BPH-trial entry criteria to the first test and dropped men who did not qualify, the second test showed improvement with no intervention: 1.0 to 1.4 points on the symptom index and 1.4 to 1.7 mL/s on peak flow, all statistically significant, and larger as the entry criteria tightened.
The lesson is not that placebo response is fake. It is that a share of any improvement in any group that was recruited because it scored badly is built in by the way it was chosen. A control arm subtracts that share out. A before-and-after cannot.
Tongkat ali trials, read from the control arm
Tongkat ali trials use different measures: hormone levels, mood scales and sexual questionnaires. The control-arm logic is the same. The existing tongkat ali post covers what the trials gave. Here the question is what the placebo and comparison groups can tell you about the reported results, using four trials.
Ismail 2012. This 12-week trial randomised 109 men aged 30 to 55 to 300 mg a day of a tongkat ali water extract or placebo (Ismail 2012). Its abstract reports better scores for the tongkat group, including “sexual libido (14% by week 12)” and sperm motility “at 44.4%”. The full paper adds context. Its own statement of the general analysis is that there were no overall significant mean differences over time between the tongkat ali and placebo groups, with significant improvements and group differences in various items in several questionnaire domains. The 14 per cent libido figure is a rise from baseline within the tongkat group on one questionnaire item, and the paper reports that the overall libido score fell in both groups between baseline and week 6 before rising in the tongkat group. The 44.4 per cent motility figure comes from a subgroup of 11 men chosen because their baseline motility was below the median; in the placebo men with low starting values, motility also rose, from 36.25 to 44.6 per cent (P = 0.252, not significant). Choosing a subgroup on its baseline value is exactly the setting Barnett describes for regression to the mean. The paper also notes that the placebo group started with a significantly higher total testosterone (18.8 against 16.5 nmol/L), a reminder that randomisation in a trial of 109 balances groups on average and not in every measure, and that scores at baseline were already about 70 per cent of the maximum, which the authors call a ceiling effect. Of the 109 randomised men, 102 were in the main analysis, 52 on tongkat ali and 50 on placebo.
Talbott 2013. This four-week trial assessed 63 moderately stressed adults, 32 men and 31 women, supplemented with a standardised hot-water tongkat ali extract or a look-alike placebo (Talbott 2013). On the mood scale, the tongkat group did better than placebo on tension (−11%), anger (−12%) and confusion (−15%), and there was no difference between groups on depression, vigor or fatigue. Vigor and fatigue are the two indices closest to what a front label calls energy, and the comparison found nothing there. Salivary testosterone was 37 per cent higher and cortisol 16 per cent lower in the tongkat group compared with placebo. The abstract shows what a control group makes possible: the reader can tell which mood measures moved and which did not.
Udani 2014. This 12-week pilot randomised 30 men to a two-extract product or placebo, with 62 men screened and four leaving early. The analysis was of completers only, 12 on the product and 14 on placebo (Udani 2014). In a trial this small, every man who leaves is more than 3 per cent of the total, and the men who stay are not a random sample of the men who started.
Henkel 2014. As its abstract describes it, this was a single group of 25 physically active seniors, 13 men and 12 women aged 57 to 72, who took 400 mg of tongkat ali extract daily for five weeks, with measurements compared before and after (Henkel 2014). No placebo group is described. It found rises in total and free testosterone and in handgrip strength. It is a fair pilot study, and a good example of the kind of result that cannot separate the extract from everything else that changed over five weeks in people who were all being given the extract.
None of that shows tongkat ali fails to work. Some of these trials do find differences between arms, and a 2022 meta-analysis of randomised trials found higher total testosterone in men given tongkat ali (Leisegang 2022). It shows that the honest reading of each of these papers is the gap between its arms, and that the arms often differ by less than the headline suggests.
Placebo effect versus placebo response
It is tempting to read a big placebo-arm gain as proof that the mind can do a lot. The best-known attempt to measure the placebo effect proper is the Cochrane review of placebo interventions for all clinical conditions (Hróbjartsson and Gøtzsche 2010). Its method is the point: it included only randomised trials that had both a placebo arm and a no-treatment arm, because that is the only design that separates the placebo from everything else that changes over time.
The review included 234 trials, 202 of them with usable outcome data, across 60 conditions. For binary outcomes the pooled relative risk for placebo against no treatment was 0.93 (95 per cent confidence interval 0.88 to 0.99), and for continuous outcomes the standardised mean difference was −0.23 (−0.28 to −0.17), larger for patient-reported outcomes (−0.26) than for observer-reported ones (−0.13). The authors called it questionable to pool the continuous trials at all, given the variation between small and large trials, and rated only 16 of the trials, 8 per cent, at low risk of bias. Read in plain terms, the placebo effect against no treatment is real but modest, and it shows up most in what people say about themselves.
That fits the pattern above, where symptom scores in placebo arms moved and flow rates did not. It also explains why a symptom questionnaire, a mood scale or a libido item is the kind of outcome where a supplement’s before-and-after is most easily inflated, and why a well-run trial is worth more there than anywhere.
The same pattern turns up outside urology. In a small crossover trial of a wild yam cream in 23 menopausal women, symptom diaries showed a minor effect of both placebo and active cream on flushing, and no significant difference between the two (Komesaroff 2001).
How to read your own before-and-after
You cannot run a trial on yourself, but a few habits make a personal before-and-after much less misleading. None of them needs equipment.
- Pick the measure before you start. If it is urinary symptoms, use the same validated questionnaire each time (the AUA Symptom Index and the IPSS are the ones in these trials). If it is energy or mood, use the same simple 0-to-10 rating at the same time of day. Deciding afterwards what counts as better is how anyone finds an improvement.
- Take more than one baseline. Score yourself on several days across a couple of weeks. In the Sech study, healthy men retested 10 to 20 days apart did not repeat their own scores exactly, and starting when you feel worst builds regression to the mean into your result.
- Know how big the noise is. The placebo arms above improved by about 3 to 4.4 symptom-score points. A gain of that size on your own score is inside the range that people taking nothing active show. A larger and lasting gain is still not proof, but a gain of a few points is not evidence at all.
- Change one thing at a time. A new supplement started in the same month as a new training plan, better sleep or less alcohol cannot be credited or blamed on its own.
- Notice which of your measures are easy to influence. In the plant-extract placebo arms, scores changed and flow rates did not. Ratings of how you feel respond to expectation. That is not a reason to distrust how you feel, only a reason to hold a rating loosely.
- Write down anything new. New or worsening symptoms, or anything unexpected, belong with a clinician, and prescribed treatment should not be stopped or changed on the strength of a supplement result.
What a personal before-and-after can tell you is whether you feel better or worse. What it cannot tell you is why. Pausing the capsule for a few weeks and watching what happens is a cheap extra check, though it is not blinded either, and it inherits every problem above.
AlphaSteel, with every amount printed on the label
Two capsules a day, 60 capsules to a bottle, seven actives with their milligrams on the panel.
Order AlphaSteelWhat this leaves standing
Nothing here says saw palmetto or tongkat ali does nothing. The large saw palmetto trials found no advantage over placebo in men with urinary symptoms, the tongkat ali trials found some between-arm differences on some measures in some groups, and no trial has tested this capsule. What the control arms add is a way of reading every one of those results: as a gap between two groups, on the measure the trial chose, in the men it studied.
The front of the AlphaSteel bottle makes support statements about energy, male vitality and confidence. The trials discussed here measured urinary scores, mood indices, questionnaires and hormone levels, and none of them measured those statements in men taking this capsule. That is not a criticism of the label, which is written to the rules for a supplement, and a small reason to weigh a before-and-after, your own or someone else’s, with the control group in mind.
AlphaSteel is a dietary supplement, not a medicine. It is not a treatment for a prostate condition, urinary symptoms, erectile dysfunction or low testosterone, and it is not a substitute for prescribed medicine. If you have urinary symptoms or a medical condition, speak to a physician; the post on the caution line goes through who the label asks to check first.
Sources behind this AlphaSteel article
- Bent S, Kane C, Shinohara K, Neuhaus J, Hudes ES, Goldberg H, et al. Saw palmetto for benign prostatic hyperplasia. N Engl J Med. 2006;354(6):557-66. doi:10.1056/NEJMoa053085 PMID 16467543. https://pubmed.ncbi.nlm.nih.gov/16467543/
- Barry MJ, Meleth S, Lee JY, Kreder KJ, Avins AL, Nickel JC, et al. Effect of increasing doses of saw palmetto extract on lower urinary tract symptoms: a randomized trial. JAMA. 2011;306(12):1344-51. doi:10.1001/jama.2011.1364 PMID 21954478. https://pubmed.ncbi.nlm.nih.gov/21954478/
- Franco JV, Trivisonno L, Sgarbossa NJ, Alvez GA, Fieiras C, Escobar Liquitay CM, et al. Serenoa repens for the treatment of lower urinary tract symptoms due to benign prostatic enlargement. Cochrane Database Syst Rev. 2023;6(6):CD001423. doi:10.1002/14651858.CD001423.pub4 PMID 37345871. https://pubmed.ncbi.nlm.nih.gov/37345871/
- Lopatkin N, Sivkov A, Walther C, Schläfke S, Medvedev A, Avdeichuk J, et al. Long-term efficacy and safety of a combination of sabal and urtica extract for lower urinary tract symptoms: a placebo-controlled, double-blind, multicenter trial. World J Urol. 2005;23(2):139-46. doi:10.1007/s00345-005-0501-9 PMID 15928959. https://pubmed.ncbi.nlm.nih.gov/15928959/
- Sökeland J, Albrecht J. [Combination of Sabal and Urtica extract vs. finasteride in benign prostatic hyperplasia (Aiken stages I to II). Comparison of therapeutic effectiveness in a one year double-blind study]. Urologe A. 1997;36(4):327-33. doi:10.1007/s001200050106 PMID 9340898. https://pubmed.ncbi.nlm.nih.gov/9340898/
- Eredics K, Madersbacher S, Schauer I. A relevant midterm (12 months) placebo effect on lower urinary tract symptoms and maximum flow rate in male lower urinary tract symptom and benign prostatic hyperplasia: a meta-analysis. Urology. 2017;106:160-6. doi:10.1016/j.urology.2017.05.011 PMID 28506862. https://pubmed.ncbi.nlm.nih.gov/28506862/
- Sech SM, Montoya JD, Bernier PA, Barnboym E, Brown S, Gregory A, et al. The so-called “placebo effect” in benign prostatic hyperplasia treatment trials represents partially a conditional regression to the mean induced by censoring. Urology. 1998;51(2):242-50. doi:10.1016/s0090-4295(97)00609-2 PMID 9495705. https://pubmed.ncbi.nlm.nih.gov/9495705/
- Barnett AG, van der Pols JC, Dobson AJ. Regression to the mean: what it is and how to deal with it. Int J Epidemiol. 2005;34(1):215-20. doi:10.1093/ije/dyh299 PMID 15333621. https://pubmed.ncbi.nlm.nih.gov/15333621/
- Hróbjartsson A, Gøtzsche PC. Placebo interventions for all clinical conditions. Cochrane Database Syst Rev. 2010;(1):CD003974. doi:10.1002/14651858.CD003974.pub3 PMID 20091554. https://pubmed.ncbi.nlm.nih.gov/20091554/
- Ismail SB, Wan Mohammad WM, George A, Nik Hussain NH, Musthapa Kamal ZM, Liske E. Randomized clinical trial on the use of PHYSTA freeze-dried water extract of Eurycoma longifolia for the improvement of quality of life and sexual well-being in men. Evid Based Complement Alternat Med. 2012;2012:429268. doi:10.1155/2012/429268 PMID 23243445. https://pubmed.ncbi.nlm.nih.gov/23243445/
- Talbott SM, Talbott JA, George A, Pugh M. Effect of tongkat ali on stress hormones and psychological mood state in moderately stressed subjects. J Int Soc Sports Nutr. 2013;10:28. doi:10.1186/1550-2783-10-28 PMID 23705671. https://pubmed.ncbi.nlm.nih.gov/23705671/
- Udani JK, George AA, Musthapa M, Pakdaman MN, Abas A. Effects of a proprietary freeze-dried water extract of Eurycoma longifolia (Physta) and Polygonum minus on sexual performance and well-being in men: a randomized, double-blind, placebo-controlled study. Evid Based Complement Alternat Med. 2014;2014:179529. doi:10.1155/2014/179529 PMID 24550993. https://pubmed.ncbi.nlm.nih.gov/24550993/
- Henkel RR, Wang R, Bassett SH, Chen T, Liu N, Zhu Y, et al. Tongkat ali as a potential herbal supplement for physically active male and female seniors: a pilot study. Phytother Res. 2014;28(4):544-50. doi:10.1002/ptr.5017 PMID 23754792. https://pubmed.ncbi.nlm.nih.gov/23754792/
- Komesaroff PA, Black CV, Cable V, Sudhir K. Effects of wild yam extract on menopausal symptoms, lipids and sex hormones in healthy menopausal women. Climacteric. 2001;4(2):144-50. PMID 11428178. https://pubmed.ncbi.nlm.nih.gov/11428178/
- Leisegang K, Finelli R, Sikka SC, Panner Selvam MK. Eurycoma longifolia (Jack) improves serum total testosterone in men: a systematic review and meta-analysis of clinical trials. Medicina (Kaunas). 2022;58(8):1047. doi:10.3390/medicina58081047 PMID 36013514. https://pubmed.ncbi.nlm.nih.gov/36013514/
- National Center for Complementary and Integrative Health. Saw palmetto. Page read September 27, 2026. https://www.nccih.nih.gov/health/saw-palmetto
AlphaSteel, with every amount printed on the label
Two capsules a day, 60 capsules to a bottle, and seven actives with their milligrams on the panel.
Order AlphaSteel