Why Most Published Research Findings Are False
2005年8月、創刊まもないオープンアクセス誌 PLoS Medicine に、この一本が載った。冒頭に置かれた掲載欄の断り書きが言うとおり、これは研究報告ではなく「一般の医学読者にとって広く関心のある主題についての意見記事」、つまりエッセイである。著者は臨床研究と疫学の方法論を専門とする研究者で、本文は、発表された知見が後続の証拠に覆されるという出来事が、臨床試験や伝統的な疫学から最先端の分子研究にいたるまで、研究デザインの全域で起きていると述べたところから始まる。そのうえで、こう言い切る——驚くにはあたらない、主張されている研究知見の大半が偽であることは、証明できるのだ、と。この挑発的な題は以後よく引かれ、いま「再現性の危機」と呼ばれている議論の出発点のひとつになった。
論証に使う道具はごく初等的である。ある科学分野で吟味される関係のうち、真の関係の数と「関係が存在しない」ものの数との比を R と置く。そこに、有意と判定する敷居 α と、真の関係を捕まえられる確率である検出力 1 − β を掛け合わせて、2×2表を作る。すると、有意と出て世に出た主張のうち実際に真であるものの割合——陽性的中率 PPV——が、R と α と β だけの式として書き下せる。論文はここに二つの現実的な要素を足す。ひとつはバイアス u、すなわち、設計・データ・解析・提示のうえでの操作によって、本来なら知見にならなかったものを知見として押し出してしまう傾向。もうひとつは、同じ問いに世界じゅうの何十というチームが取り組み、そのうち一つでも有意を出せばそれが注目を集める、という事情である。これらを動かすと PPV がどう動くかを見て、そこから「系」を番号つきで並べていく。これが全体の運びである。
読む前に三つ。第一に、この論文の「研究知見」と「偽」は、いずれも狭く定義されている。研究知見とは、統計的有意性に達した任意の関係のこと——有効な介入、情報量のある予測因子、危険因子、連関——であり、偽とは、その関係が実際には存在しないという意味である。捏造や計算間違いのことではない。第二に、この論証は条件つきである。PPV は R・α・β・u・チーム数 n の関数であって、これらの値そのものを論文が測ったわけではない。分野ごとにもっともらしい値を置いて、そのときどうなるかを見せている。だから読みどころは「大半は偽だ」という結論そのものよりも、どの値をどこに置けば数字がどこまで崩れるか、という感度のほうにある。第三に、この枠組みは「p < 0.05 で二分する」という慣行を前提として組まれている。論文はまさにその慣行を批判しているのだが、批判の足場としてはそれを使う。批判の対象と道具が同じものだという、ややねじれた構えである。
全2回の第1回であるこの回は、要旨と導入、2×2表による PPV の導出、バイアスと複数チームのモデル、全ゲノム関連研究を例にとった具体例、そして系1から系3までを読む。系1〜3は、研究の規模・効果量・検定される関係の数という、いわば研究の「かたち」にかかわる三つである。なお、掲載誌のページには2022年8月付の訂正へのリンクが置かれている(本文の一行目に見えるのがそれである)。
この論文が使う数学は、掛け算と割り算、それに確率の初歩だけである。ただ、術語が立て続けに出てくるので、先に道具を四つ並べておく。以下に出す数値例は、説明のために編者が作ったもので、論文にある数字ではない。
道具1 α と 1 − β——二種類の間違え方
検定は、「関係なし」という想定を立てて、手元のデータがその想定のもとでどれだけ珍しいかを見る。珍しさの敷居を有意水準 α と呼び、慣例で 0.05 に置く。本当は関係がないのに「ある」と言ってしまうことを二十回に一回まで許す、という取り決めである。これが第一種の過誤。逆に、本当は関係があるのに取り逃がすのが第二種の過誤で、その率を β と書き、1 − β を検出力という。関係が本当にあるとき、それを捕まえられる確率である。煙感知器にたとえるなら、α は料理の湯気で鳴ってしまう率、1 − β は本物の火事でちゃんと鳴る率にあたる。ここで大事なのは、α が慣例で固定されているのに対し、検出力のほうは標本の大きさと、探している効果の大きさで決まってしまう、ということだ。論文の系1と系2は、まさにこの一点から出てくる。
道具2 研究前オッズ R
ある分野で吟味される関係のうち、真の関係の数と「関係なし」の数との比を R と置く。比が 1 対 9 なら R = 1/9 で、任意の一つが真である確率は R/(R + 1) = 1/10。比を確率に直すだけの話である。R は分野に固有で、桁で違いうる。ありそうな度合いの高い関係を狙い撃ちにする分野と、措定されうる何百万もの仮説のなかからただ一つを探す分野とでは、まるで違う。編者の例で言えば、候補を絞りきった確証的な試験なら真と偽が 1 対 1 に近いこともあろうし、総当たりの探索なら 1 対 10万ということもある。この一つの数が、以下のすべてを支配する。
道具3 2×2表から PPV へ
c 個の関係を検定するとしよう。そのうち真の関係は c・R/(R + 1) 個、関係なしは c・1/(R + 1) 個ある。真の関係のうち有意になるのは検出力 1 − β の割合、関係なしのうち有意になってしまうのは α の割合。だから「有意になった」もの全体は (1 − β)・cR/(R + 1) と α・c/(R + 1) の和で、このうち真であるものの割合が PPV である。c も (R + 1) も約分されて、
PPV = (1 − β)R / ((1 − β)R + α) = (1 − β)R / (R − βR + α)
これが本文に出てくる式である。目を留めるべきは、分母に α がそのまま残ることだ。真の側は R によって薄められるのに、偽陽性の側は薄まらない。R が小さい分野——つまり探索的な分野——では、(1 − β)R が α に負ける。そうなった瞬間、PPV は 0.5 を割る。論文が「(1 − β)R > α であれば、その知見は偽であるよりも真である見込みのほうが大きい」と書くのは、この分かれ目のことである。数を入れてみると、この分かれ目がいかに近いかが分かる。
十本に一本が真という、それほど悲観的でもない分野で、検出力が半分。それだけで、発表された知見の当たり外れはほぼ五分五分になる。検出力が 0.2 に落ちれば三割を切る。ここまでが、バイアスも競争もない、誰もが正直に振る舞った場合の話である。
道具4 バイアス u と、チームの数 n
u は、探索された解析のうち、本来なら「研究知見」にならなかったはずなのに、操作のせいで結局そう報告されてしまうものの割合である。患者の都合のよい出し入れ、事後のサブグループ解析、当初指定していなかった比較の追加、定義の変更、選択的な報告——本文が挙げるのはそうしたものだ。u が加わると、真の関係のうち取り逃していたものの一部も拾い上げられるので分子も少し増えるが、関係なしの側からはそれ以上に多くが繰り上がる。差し引きで PPV は下がる。バイアスを偶然変動と混同しないこと、と本文がわざわざ念を押すのは、この u が「完璧に行っても残る運の悪さ」とは別建ての量だからである。
もう一つが n、同じ問いに取り組む独立なチームの数。一つでも有意を出せばそれが知見として世に出る、という状況を考える。関係が本当にないとき、n 回のうち少なくとも一回まちがって有意になる確率は 1 − (1 − α)ⁿ である。α = 0.05 なら、n = 14 で五割を超える(0.95¹⁴ ≈ 0.49。編者の計算)。一方、真の関係を少なくとも一つのチームが捕まえる確率は 1 − βⁿ で、こちらも n とともに増えるが、上限の 1 に早く頭打ちになる。両方が増えるので、正味でどちらに転ぶかは目分量では決まらない。論文はこれを同じ2×2表に入れて計算する。その帰結が何であるかは、次回の系にかかわる。
この回の鍵語はほとんどが統計学の術語で、日本語の定訳が分野によって揺れるものが多い。本文の訳では次の語をあてた。
| 原語 | 読み | 訳と意味 |
| research finding | リサーチ・ファインディング | 研究知見。この論文の定義では「統計的有意性に達した任意の関係」——有効な介入、情報量のある予測因子、危険因子、連関など。日常語の「発見」より狭く、機械的である。 |
| false | フォールス | 偽。主張された関係が実際には存在しない、という意味。捏造や計算間違いのことではない。「誤り」と訳すと後者に寄るので「偽」を採った。 |
| power (1 − β) | パワー | 検出力。検定力とも訳される。真の関係があるとき、それを有意と判定できる確率。標本が大きいほど、また効果が大きいほど高い。 |
| positive predictive value (PPV) | ポジティヴ・プレディクティヴ・ヴァリュー | 陽性的中率。有意と主張された知見のうち、実際に真であるものの割合。診断検査の語を、研究そのものの当たり外れに転用したもの。 |
| pre-study odds (R) | プリスタディ・オッズ | 研究前オッズ。ある分野で吟味される関係のうち、真の関係の数と「関係なし」の数との比。確率に直せば R/(R + 1)。分野に固有の値で、桁で違いうる。 |
| bias (u) | バイアス | バイアス。本来なら知見にならないものを知見として生み出す、設計・データ・解析・提示上の要因の総体。完璧な研究でも起きる偶然変動とは、はっきり別物とされる。 |
| effect size | エフェクト・サイズ | 効果量。関係の大きさ。相対リスクやオッズ比で測る。喫煙と癌なら3〜20、多遺伝子性疾患の遺伝的危険因子なら1.1〜1.5、というように分野で桁が違う。 |
| null | ヌル | 「関係なし」。真の関係が存在しない状態。統計学では「帰無」の語をあてる場面もあるが、本文では読みやすさを採って「関係なし」とした。 |
| corollary | コロラリー | 系。証明された事柄から直ちに導かれる帰結。この論文の後半は「系1、系2……」と番号を振って並ぶ。 |
| gold standard | ゴールド・スタンダード | 金標準。突き合わせの基準となる真の姿。ここでは「その関係が実際に存在するかどうか」そのものを指す。 |
Essays are opinion pieces on a topic of broad interest to a general medical audience. Aug 2022: Ioannidis JPA (2022) Correction: Why Most Published Research Findings Are False. PLOS Medicine 19(8): e1004085. https://doi.org/10.1371/journal.pmed.1004085 View correction
エッセイとは、一般の医学読者にとって広く関心のある主題についての意見記事である。 2022年8月:Ioannidis JPA (2022)「訂正:なぜ発表される研究知見の大半は偽なのか」PLOS Medicine 19(8): e1004085. https://doi.org/10.1371/journal.pmed.1004085 訂正を見る
There is increasing concern that most current published research findings are false. The probability that a research claim is true may depend on study power and bias, the number of other studies on the same question, and, importantly, the ratio of true to no relationships among the relationships probed in each scientific field. In this framework, a research finding is less likely to be true when the studies conducted in a field are smaller; when effect sizes are smaller; when there is a greater number and lesser preselection of tested relationships; where there is greater flexibility in designs, definitions, outcomes, and analytical modes; when there is greater financial and other interest and prejudice; and when more teams are involved in a scientific field in chase of statistical significance. Simulations show that for most study designs and settings, it is more likely for a research claim to be false than true. Moreover, for many current scientific fields, claimed research findings may often be simply accurate measures of the prevailing bias. In this essay, I discuss the implications of these problems for the conduct and interpretation of research.
現在発表されている研究知見の大半は偽である、という懸念が高まっている。ある研究上の主張が真である確率は、研究の検定力とバイアス、同じ問いをめぐる他の研究の数、そして重要なことに、それぞれの科学分野で吟味される関係のうち、真の関係と、関係が存在しないものとの比に依存しうる。この枠組みでは、ある研究知見が真である見込みは次のような場合に低くなる——その分野で行われる研究の規模がより小さいとき、効果量がより小さいとき、検証される関係の数がより多く、その事前選別がより乏しいとき、デザイン・定義・アウトカム・解析法における自由度がより大きいとき、金銭上のものその他の利害と予断がより大きいとき、そして統計的有意性を追ってその科学分野により多くの研究チームが参入しているとき。シミュレーションが示すところでは、大半の研究デザインと設定において、研究上の主張は真であるよりも偽である見込みのほうが高い。さらに、現在の多くの科学分野において、主張される研究知見は、しばしばその分野に行き渡っているバイアスを正確に測った値にすぎない、ということがありうる。本エッセイでは、これらの問題が研究の遂行と解釈に対して何を含意するかを論じる。
Citation: Ioannidis JPA (2005) Why Most Published Research Findings Are False. PLoS Med 2(8): e124. https://doi.org/10.1371/journal.pmed.0020124
引用情報:Ioannidis JPA (2005)「なぜ発表される研究知見の大半は偽なのか」PLoS Med 2(8): e124. https://doi.org/10.1371/journal.pmed.0020124
Copyright: © 2005 John P. A. Ioannidis. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
著作権:© 2005 John P. A. Ioannidis。本稿は、クリエイティブ・コモンズ表示ライセンスの条件のもとで配布されるオープンアクセス論文である。原著作が適切に引用されるかぎり、いかなる媒体においても、無制限の使用・配布・複製が許される。
Competing interests: The author has declared that no competing interests exist. Abbreviation: PPV, positive predictive value
競合する利害〔の開示〕:著者は、競合する利害は存在しないと表明した。略号:PPV、陽性的中率
Published research findings are sometimes refuted by subsequent evidence, with ensuing confusion and disappointment. Refutation and controversy is seen across the range of research designs, from clinical trials and traditional epidemiological studies [1–3] to the most modern molecular research [4,5]. There is increasing concern that in modern research, false findings may be the majority or even the vast majority of published research claims [6–8]. However, this should not be surprising. It can be proven that most claimed research findings are false. Here I will examine the key factors that influence this problem and some corollaries thereof.
発表された研究知見は、後続の証拠によって覆されることがあり、そのたびに混乱と失望が生じる。反証と論争は、臨床試験や伝統的な疫学研究[1–3]から、最も現代的な分子研究[4,5]にいたるまで、研究デザインの全域にわたって見られる。現代の研究においては、偽なる知見が、発表される研究上の主張の過半数、いや圧倒的多数を占めるのではないか——そうした懸念が高まっている[6–8]。しかし、これは驚くにはあたらない。主張されている研究知見の大半が偽であることは、証明できるのである。ここで私は、この問題を左右する主要な因子と、そこから導かれるいくつかの系とを検討する。
Several methodologists have pointed out [9–11] that the high rate of nonreplication (lack of confirmation) of research discoveries is a consequence of the convenient, yet ill-founded strategy of claiming conclusive research findings solely on the basis of a single study assessed by formal statistical significance, typically for a p-value less than 0.05. Research is not most appropriately represented and summarized by p-values, but, unfortunately, there is a widespread notion that medical research articles should be interpreted based only on p-values. Research findings are defined here as any relationship reaching formal statistical significance, e.g., effective interventions, informative predictors, risk factors, or associations. “Negative” research is also very useful. “Negative” is actually a misnomer, and the misinterpretation is widespread. However, here we will target relationships that investigators claim exist, rather than null findings.
何人かの方法論家が指摘している[9–11]——研究上の発見が再現されない(確証が得られない)割合の高さは、便利ではあるが根拠の乏しいひとつの戦略の帰結である、と。すなわち、形式的な統計的有意性——典型的にはp値が0.05未満であること——によって評価された単一の研究だけを根拠に、決定的な研究知見を主張するという戦略の帰結である。研究は、p値によって表現され要約されるのが最も適切だというわけではない。ところが残念なことに、医学研究の論文はp値のみに基づいて解釈されるべきだという通念が広く行きわたっている。ここでは「研究知見」を、形式的な統計的有意性に達した任意の関係と定義する。たとえば、有効な介入、情報量のある予測因子、危険因子、あるいは連関である。「陰性」の研究もまた、きわめて有用である。「陰性」というのは実のところ呼び名として適切でなく、その誤った解釈は広く行きわたっている。しかしここでは、帰無的な知見ではなく、研究者が存在すると主張する関係を対象とする。
It can be proven that most claimed research findings are false As has been shown previously, the probability that a research finding is indeed true depends on the prior probability of it being true (before doing the study), the statistical power of the study, and the level of statistical significance [10,11]. Consider a 2 × 2 table in which research findings are compared against the gold standard of true relationships in a scientific field. In a research field both true and false hypotheses can be made about the presence of relationships. Let R be the ratio of the number of “true relationships” to “no relationships” among those tested in the field. R is characteristic of the field and can vary a lot depending on whether the field targets highly likely relationships or searches for only one or a few true relationships among thousands and millions of hypotheses that may be postulated. Let us also consider, for computational simplicity, circumscribed fields where either there is only one true relationship (among many that can be hypothesized) or the power is similar to find any of the several existing true relationships.
主張された研究知見の大半が偽であることは証明できる
すでに示されているとおり、ある研究知見が実際に真である確率は、それが真であることの事前確率(研究を行う前の)と、その研究の統計的検出力と、統計的有意性の水準とに依存する[10,11]。ある科学分野において、研究知見を、真の関係という金標準と突き合わせる2×2表を考えてみよう。一つの研究分野では、関係の存在について、真の仮説も偽の仮説も立てられうる。その分野で検定されるもののうち、「真の関係」の数と「関係なし」の数との比を R としよう。R はその分野に固有の値であり、大きく変わりうる——その分野が、ありそうな度合いの高い関係を狙うのか、それとも措定されうる何千何百万もの仮説のなかから、ただ一つないしごく少数の真の関係を探すのかによって。さらに、計算を簡単にするため、限定された分野を考えることにしよう。すなわち、(仮説として立てうる多くのもののうち)真の関係がただ一つしか存在しないか、あるいは、現に存在する複数の真の関係のどれを見出す場合でも検出力が同程度であるような分野である。
The pre-study probability of a relationship being true is R/(R + 1). The probability of a study finding a true relationship reflects the power 1 - β (one minus the Type II error rate). The probability of claiming a relationship when none truly exists reflects the Type I error rate, α. Assuming that c relationships are being probed in the field, the expected values of the 2 × 2 table are given in Table 1. After a research finding has been claimed based on achieving formal statistical significance, the post-study probability that it is true is the positive predictive value, PPV. The PPV is also the complementary probability of what Wacholder et al. have called the false positive report probability . According to the 2 × 2 table, one gets PPV = (1 - β)R/(R - βR + α). A research finding is thus more likely true than false if (1 - β)R > α. Since usually the vast majority of investigators depend on a = 0.05, this means that a research finding is more likely true than false if (1 - β)R > 0.05.
ある関係が真であるという研究前確率は R/(R + 1) である。ある研究が真の関係を見出す確率は、検出力 1 − β[1から第二種の過誤の率を引いたもの]を反映する。真には関係が存在しないのに関係を主張してしまう確率は、第一種の過誤の率 α を反映する。その分野で c 個の関係が探索されていると仮定すると、2×2表の期待値は表1に示すとおりとなる。形式的な統計的有意性を達成したことを根拠にある研究知見が主張されたのち、それが真であるという研究後確率が、陽性的中率 PPV である。PPV はまた、Wacholder らが偽陽性報告確率と呼んだものの、余事象の確率でもある。2×2表によれば、PPV = (1 − β)R/(R − βR + α) が得られる。したがって、(1 − β)R > α であれば、その研究知見は偽であるよりも真である見込みのほうが大きい。研究者の圧倒的多数は通常 α = 0.05に依拠しているのだから、これはつまり、(1 − β)R > 0.05 であれば研究知見は偽であるより真である見込みが大きい、ということを意味する。
https://doi.org/10.1371/journal.pmed.0020124.t001 What is less well appreciated is that bias and the extent of repeated independent testing by different teams of investigators around the globe may further distort this picture and may lead to even smaller probabilities of the research findings being indeed true. We will try to model these two factors in the context of similar 2 × 2 tables.
〔表1〕https://doi.org/10.1371/journal.pmed.0020124.t001
あまり十分には認識されていないのは、次のことである。バイアスと、世界じゅうの異なる研究者チームによって独立した検定がどれだけ反復されるかということ、この二つが〔いま描いた〕この描像をさらに歪め、研究知見が実際に真である確率をいっそう小さくしてしまいうる、ということである。われわれは、これら二つの要因を、同様の2×2表という枠組みのなかでモデル化することを試みる。
First, let us define bias as the combination of various design, data, analysis, and presentation factors that tend to produce research findings when they should not be produced. Let u be the proportion of probed analyses that would not have been “research findings,” but nevertheless end up presented and reported as such, because of bias. Bias should not be confused with chance variability that causes some findings to be false by chance even though the study design, data, analysis, and presentation are perfect. Bias can entail manipulation in the analysis or reporting of findings. Selective or distorted reporting is a typical form of such bias. We may assume that u does not depend on whether a true relationship exists or not. This is not an unreasonable assumption, since typically it is impossible to know which relationships are indeed true. In the presence of bias (Table 2), one gets PPV = ([1 - β]R + uβR)/(R + α − βR + u − uα + uβR), and PPV decreases with increasing u, unless 1 − β ≤ α, i.e., 1 − β ≤ 0.05 for most situations. Thus, with increasing bias, the chances that a research finding is true diminish considerably.
まず、バイアスを次のように定義しよう。すなわち、本来ならば研究知見を生み出すべきでないときにそれを生み出す傾向をもつ、設計・データ・解析・提示上のさまざまな要因の組み合わせである。探索された解析のうち、本来ならば「研究知見」とはならなかったはずなのに、それでもバイアスのゆえに結局そのようなものとして提示され報告されてしまうものの割合を u としよう。バイアスは、偶然による変動と混同されてはならない。偶然変動とは、研究の設計・データ・解析・提示が完璧であってもなお、いくつかの知見をたまたま偽にしてしまうもののことである。バイアスは、解析における操作や、知見の報告における操作を伴いうる。選択的な報告、あるいは歪められた報告は、そうしたバイアスの典型的な形である。u は、真の関係が存在するかどうかには依存しない、と仮定してよいだろう。これは不合理な仮定ではない。というのも、どの関係が実際に真であるのかを知ることは、通常は不可能だからである。バイアスが存在する場合(表2)、PPV = ([1 − β]R + uβR)/(R + α − βR + u − uα + uβR) が得られ、PPV は u の増加とともに減少する。ただし 1 − β ≤ α である場合、すなわちたいていの状況では 1 − β ≤ 0.05 である場合を除く。したがって、バイアスが増すにつれて、ある研究知見が真である見込みはかなりの程度まで小さくなる。
This is shown for different levels of power and for different pre-study odds in Figure 1. Conversely, true research findings may occasionally be annulled because of reverse bias. For example, with large measurement errors relationships are lost in noise , or investigators use data inefficiently or fail to notice statistically significant relationships, or there may be conflicts of interest that tend to “bury” significant findings . There is no good large-scale empirical evidence on how frequently such reverse bias may occur across diverse research fields. However, it is probably fair to say that reverse bias is not as common. Moreover measurement errors and inefficient use of data are probably becoming less frequent problems, since measurement error has decreased with technological advances in the molecular era and investigators are becoming increasingly sophisticated about their data. Regardless, reverse bias may be modeled in the same way as bias above. Also reverse bias should not be confused with chance variability that may lead to missing a true relationship because of chance.
このことは、さまざまな水準の検出力について、またさまざまな研究前オッズについて、図1に示されている。逆に、真の研究知見が、逆バイアスのゆえに取り消されてしまうことも、ときにはありうる。たとえば、測定誤差が大きければ関係は雑音のなかに埋もれて失われるし、あるいは研究者がデータを非効率的に用いたり、統計的に有意な関係を見落としたりすることもあるし、あるいはまた、有意な知見を「葬り去る」傾向をもつ利益相反が存在することもある。多様な研究分野にわたってそのような逆バイアスがどれほどの頻度で生じうるのかについて、大規模で良質な経験的証拠は存在しない。しかしおそらく、逆バイアスはそれほど一般的ではない、と言ってよいだろう。さらに、測定誤差やデータの非効率的な使用は、おそらくそれほど頻繁な問題ではなくなりつつある。というのも、分子〔生物学〕の時代における技術の進歩とともに測定誤差は減少してきたし、研究者たちは自分のデータについてますます洗練されてきているからである。いずれにせよ、逆バイアスは、上述のバイアスと同じやり方でモデル化しうる。また、逆バイアスは、偶然のために真の関係を見落とすことにつながりうる偶然変動と混同されてはならない。
Panels correspond to power of 0.20, 0.50, and 0.80. https://doi.org/10.1371/journal.pmed.0020124.g001 https://doi.org/10.1371/journal.pmed.0020124.t002
各パネルは、検出力0.20、0.50、0.80に対応する。 https://doi.org/10.1371/journal.pmed.0020124.g001 https://doi.org/10.1371/journal.pmed.0020124.t002
Several independent teams may be addressing the same sets of research questions. As research efforts are globalized, it is practically the rule that several research teams, often dozens of them, may probe the same or similar questions. Unfortunately, in some areas, the prevailing mentality until now has been to focus on isolated discoveries by single teams and interpret research experiments in isolation. An increasing number of questions have at least one study claiming a research finding, and this receives unilateral attention. The probability that at least one study, among several done on the same question, claims a statistically significant research finding is easy to estimate. For n independent studies of equal power, the 2 × 2 table is shown in Table 3: PPV = R(1 − β)/(R + 1 − [1 − α] − Rβ) (not considering bias). With increasing number of independent studies, PPV tends to decrease, unless 1 - β < a, i.e., typically 1 − β < 0.05. This is shown for different levels of power and for different pre-study odds in Figure 2. For n studies of different power, the term β is replaced by the product of the terms βi for i = 1 to n, but inferences are similar.
複数の独立した研究チームが、同じ研究上の問いの組に取り組んでいることがありうる。研究の営みが世界規模になるにつれ、複数の研究チーム——しばしば数十のチーム——が同一の、あるいは類似の問いを探るということは、事実上の通例となっている。残念なことに、いくつかの分野では、これまで支配的であった心性は、単一のチームによる孤立した発見に注目し、研究上の実験を孤立させたまま解釈するというものであった。少なくとも一つの研究が研究知見を主張しているような問いはますます数を増しており、そしてその〔一つの研究〕が一方的に注目を集める。同じ問いについて行われた複数の研究のうち、少なくとも一つが統計的に有意な研究知見を主張する確率は、容易に見積もることができる。検出力の等しいn個の独立な研究については、2×2表は表3に示したとおりであって、PPV = R(1 − β)/(R + 1 − [1 − α] − Rβ)(バイアスは考慮していない)となる。独立した研究の数が増えるにつれて、PPVは減少する傾向にある。ただし 1 − β < α である場合、すなわち通常は 1 − β < 0.05 である場合は別である。このことは、さまざまな水準の検出力について、またさまざまな研究前オッズについて、図2に示されている。検出力の異なるn個の研究については、項βは、i = 1 から n までの項βiの積に置き換えられるが、導かれる帰結は同様である。
Panels correspond to power of 0.20, 0.50, and 0.80. https://doi.org/10.1371/journal.pmed.0020124.g002 https://doi.org/10.1371/journal.pmed.0020124.t003
各パネルは、検出力0.20、0.50、0.80に対応する。 https://doi.org/10.1371/journal.pmed.0020124.g002 https://doi.org/10.1371/journal.pmed.0020124.t003
A practical example is shown in Box 1. Based on the above considerations, one may deduce several interesting corollaries about the probability that a research finding is indeed true.
実際的な例をBox 1に示す。以上の考察にもとづいて、ある研究知見が実際に真である確率について、いくつか興味深い系を導くことができる。
Let us assume that a team of investigators performs a whole genome association study to test whether any of 100,000 gene polymorphisms are associated with susceptibility to schizophrenia. Based on what we know about the extent of heritability of the disease, it is reasonable to expect that probably around ten gene polymorphisms among those tested would be truly associated with schizophrenia, with relatively similar odds ratios around 1.3 for the ten or so polymorphisms and with a fairly similar power to identify any of them. Then R = 10/100,000 = 10, and the pre-study probability for any polymorphism to be associated with schizophrenia is also R/(R + 1) = 10. Let us also suppose that the study has 60% power to find an association with an odds ratio of 1.3 at α = 0.05. Then it can be estimated that if a statistically significant association is found with the p-value barely crossing the 0.05 threshold, the post-study probability that this is true increases about 12-fold compared with the pre-study probability, but it is still only 12 × 10.
ある研究者チームが全ゲノム関連研究を行い、10万個の遺伝子多型のいずれかが統合失調症へのかかりやすさと関連するかどうかを検定する、と仮定しよう。この疾患の遺伝率の程度について我々が知っていることにもとづけば、検定された多型のうちおよそ10個ほどが統合失調症と真に関連しており、その10個ほどの多型についてオッズ比は1.3前後で互いに比較的似通っており、そのいずれかを見つけ出す検出力もかなり似通っている、と見込むのは理にかなっている。すると R = 10/100,000 = 10〔の−4乗〕であり、任意の多型が統合失調症と関連しているという研究前確率もまた R/(R + 1) = 10〔の−4乗〕である。さらに、この研究が α = 0.05 においてオッズ比1.3の関連を見出す検出力を60%もつ、と仮定しよう。そのとき次のように見積もることができる。p値が0.05という閾値をかろうじて越える形で統計的に有意な関連が見出されたとすると、それが真である研究後確率は研究前確率と比べておよそ12倍に増える。しかしそれでもなお 12 × 10〔の−4乗〕にすぎない。
Now let us suppose that the investigators manipulate their design, analyses, and reporting so as to make more relationships cross the p = 0.05 threshold even though this would not have been crossed with a perfectly adhered to design and analysis and with perfect comprehensive reporting of the results, strictly according to the original study plan. Such manipulation could be done, for example, with serendipitous inclusion or exclusion of certain patients or controls, post hoc subgroup analyses, investigation of genetic contrasts that were not originally specified, changes in the disease or control definitions, and various combinations of selective or distorted reporting of the results. Commercially available “data mining” packages actually are proud of their ability to yield statistically significant results through data dredging. In the presence of bias with u = 0.10, the post-study probability that a research finding is true is only 4.4 × 10. Furthermore, even in the absence of any bias, when ten independent research teams perform similar experiments around the world, if one of them finds a formally statistically significant association, the probability that the research finding is true is only 1.5 × 10, hardly any higher than the probability we had before any of this extensive research was undertaken!
ここで、研究者たちが自分たちの設計・解析・報告に手を加え、より多くの関係が p = 0.05 の閾値を越えるようにする、と考えてみよう。もとの研究計画に厳密にしたがって、設計と解析を完璧に守り、結果を完璧に包括的に報告していたなら、その閾値は越えられなかったはずであるにもかかわらず、である。そうした操作は、たとえば次のようにして行いうる。ある患者や対照者を渡りに舟とばかりに組み入れたり除外したりすること、事後のサブグループ解析、当初は指定されていなかった遺伝的対比の検討、疾患や対照の定義の変更、そして結果の選択的な、あるいは歪められた報告のさまざまな組み合わせ。市販の「データマイニング」パッケージは、実のところ、データを浚うことによって統計的に有意な結果を生み出せるという能力を誇りにしている。u = 0.10 のバイアスがあるとき、ある研究知見が真である研究後確率はわずか 4.4 × 10〔の−4乗〕である。さらに、いかなるバイアスもない場合ですら、世界中で10の独立した研究チームが似たような実験を行い、そのうちの一つが形式上統計的に有意な関連を見出したとすると、その研究知見が真である確率はわずか 1.5 × 10〔の−4乗〕にすぎず、この大がかりな研究のどれ一つ着手される前に我々がもっていた確率と比べて、ほとんど高くなっていないのである!
Corollary 1: The smaller the studies conducted in a scientific field, the less likely the research findings are to be true. Small sample size means smaller power and, for all functions above, the PPV for a true research finding decreases as power decreases towards 1 − β = 0.05. Thus, other factors being equal, research findings are more likely true in scientific fields that undertake large studies, such as randomized controlled trials in cardiology (several thousand subjects randomized) than in scientific fields with small studies, such as most research of molecular predictors (sample sizes 100-fold smaller) .
系1——ある科学分野で行われる研究が小規模であるほど、その研究知見は真でありにくくなる。標本サイズが小さいということは検出力が小さいということであり、上に挙げたすべての関数について、真の研究知見に対する PPV は、検出力が 1 − β = 0.05 へと下がっていくにつれて減少する。したがって、他の要因が等しいならば、研究知見は、循環器学における無作為化比較試験(数千人の被験者が無作為に割り付けられる)のように大規模な研究を行う科学分野においてのほうが、分子予測因子の研究の大半(標本サイズは100分の1)のように小規模な研究しかない分野においてよりも、真でありやすい。
Corollary 2: The smaller the effect sizes in a scientific field, the less likely the research findings are to be true. Power is also related to the effect size. Thus research findings are more likely true in scientific fields with large effects, such as the impact of smoking on cancer or cardiovascular disease (relative risks 3–20), than in scientific fields where postulated effects are small, such as genetic risk factors for multigenetic diseases (relative risks 1.1–1.5) . Modern epidemiology is increasingly obliged to target smaller effect sizes . Consequently, the proportion of true research findings is expected to decrease. In the same line of thinking, if the true effect sizes are very small in a scientific field, this field is likely to be plagued by almost ubiquitous false positive claims. For example, if the majority of true genetic or nutritional determinants of complex diseases confer relative risks less than 1.05, genetic or nutritional epidemiology would be largely utopian endeavors.
系2——ある科学分野における効果量が小さいほど、その研究知見は真でありにくくなる。検出力は効果量とも関係している。したがって研究知見は、喫煙が癌や心血管疾患に及ぼす影響(相対リスク3〜20)のように効果の大きい科学分野においてのほうが、多遺伝子性疾患の遺伝的危険因子(相対リスク1.1〜1.5)のように想定される効果が小さい科学分野においてよりも、真でありやすい。現代の疫学は、ますます小さな効果量を標的とすることを余儀なくされている。その結果として、真である研究知見の割合は減少すると予想される。同じ筋道で考えれば、ある科学分野において真の効果量がきわめて小さい場合、その分野はほとんど至るところに偽陽性の主張がはびこるという事態に悩まされる公算が大きい。たとえば、複雑疾患の真の遺伝的あるいは栄養的な規定因子の大半が 1.05 未満の相対リスクしかもたらさないのだとすれば、遺伝疫学や栄養疫学は大部分がユートピア的な営みということになるだろう。
Corollary 3: The greater the number and the lesser the selection of tested relationships in a scientific field, the less likely the research findings are to be true. As shown above, the post-study probability that a finding is true (PPV) depends a lot on the pre-study odds (R). Thus, research findings are more likely true in confirmatory designs, such as large phase III randomized controlled trials, or meta-analyses thereof, than in hypothesis-generating experiments. Fields considered highly informative and creative given the wealth of the assembled and tested information, such as microarrays and other high-throughput discovery-oriented research [4,8,17], should have extremely low PPV.
系3——ある科学分野において検定される関係の数が多く、その選別が乏しいほど、その研究知見は真でありにくくなる。上に示したとおり、ある知見が真である研究後確率(PPV)は、研究前オッズ (R) に大きく依存する。したがって研究知見は、大規模な第III相無作為化比較試験やそのメタ解析のような確証的な設計においてのほうが、仮説生成的な実験においてよりも、真でありやすい。集積され検定される情報の豊富さのゆえに、きわめて情報量に富み創造的だと見なされている分野、たとえばマイクロアレイやその他のハイスループットな発見志向の研究 [4,8,17] は、PPV がきわめて低いはずである。
受け継がれた点。「p < 0.05 の単発の研究では足りない」という言い分自体は、2005年の時点でも新しくはなかった。本文がそれを何人かの方法論家の先行指摘に帰しているとおりである。この論文が変えたのは、その言い分を数式の形で、しかも分野をまたいで比較できる形で置き直したことだった。以後、事前オッズ・検出力・バイアスという三点セットは、研究の質を語るときの共通語彙になった。心理学での大規模な再現検証(Open Science Collaboration 2015)、前臨床がん生物学での再現失敗の報告(Begley & Ellis 2012)、事前登録と登録報告という制度、有意水準を0.005に引き下げよという提案(Benjamin ら 2018)、米国統計学会のp値に関する声明(2016)——いずれもこの見取り図の上に立っている。この回の範囲でいえば、系1(小規模な研究ほど危うい)と系2(効果量が小さい分野ほど危うい)は、検出力の低さを主要な病因とみなす今日の議論にほぼそのまま生きている。系3の理屈は、遺伝統計学が全ゲノム関連解析の有意水準を 5×10⁻⁸ という水準まで引き下げ、検定される関係の数を正面から勘定に入れるようになったことで、実務に組み込まれた。本文が例に挙げた10万個の多型という設定は、今日から見れば、その転換の直前の光景である。
否定・修正された点。 発表の二年後、グッドマンとグリーンランドが同じ誌上で、まさにこの回で扱った範囲そのものを批判した(Goodman & Greenland, “Why Most Published Research Findings Are False: Problems in the Analysis”, PLoS Medicine 2007)。要点は三つある。バイアス u の入れ方に数学上の難があること。α という敷居だけを使い、実際に観測された p 値を捨ててしまうと情報が失われること——p = 0.001 の知見と p = 0.049 の知見を同じ「陽性」として扱ってよいのか。そして、「大半は偽である」という結論は、R のような測りようのない量への仮定の上に載っており、証明されたとは言えないこと。三つ目は、この回のはじめに述べた「条件つきの論証」という性格の裏返しでもある。のちに Jager と Leek(2014)が、医学文献の要旨に報告された p 値から偽発見率をおよそ14%と見積もり、少なくとも一部の分野では「大半」ではないと論じた。この推定もまた激しく争われ、決着はついていない。もう一つの根本的な批判は、この枠組みが「関係なし」をぴったりゼロの効果として扱う点にある。効果がほぼゼロだが厳密にゼロではない関係ばかりの分野では、真と偽の二分そのものが問いを歪める、という指摘である。なお、本文の一行目にあるとおり、この論文自体、2022年に訂正が出ている。
次回は、系3の先——残る系と、この枠組みから著者が引き出す結論を読む。
底本。汎用HTML(journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.0020124)から本文を機械的に取得し、そのまま採用しました。段落分割は底本に従います。
別の転写との照合。汎用HTML の転写(pmc.ncbi.nlm.nih.gov/articles/PMC1182327/)と突き合わせたところ、約物の違いを均した本文の一致度は 95.2% でした。残る差は転写者ごとの段落の切り方と正書法の判断によるものです。