產生器 · 2026-09-22
FDA Awards $1.29 Million Contract to Test Generative AI Evaluation of Radiology Reports
重點結論
2026 年 9 月 22 日,媒體報導指出,美國食品藥物管理局(FDA)授予 Cognita Imaging 一份合約,評估生成式人工智慧在審核 AI 放射線報告中所扮演的角色,相關報導中提及的合約金額分別為 120 萬美元與 129 萬美元。此次驗證工作的目標是針對一百萬份放射線報告。該獎勵案顯示出聯邦政府正積極推動,在 AI 生成之醫療影像發現獲得核可之前,建立一套統一的評分方式。
一句話總結:值得關注的工具:具備出處標記的合成放射線報告產生器、以大型語言模型作為評審團的評分範本建構工具、模擬 DICOM 中繼資料產生器、磁振造影導引之正子造影工作流程模擬資料集、可直接提交給監管機關的評估紀錄格式化工具。
來源報導了什麼
FDA contract puts an LLM-as-a-Jury evaluation on the regulatory bench
Cognita Imaging has been awarded an FDA contract to test whether generative AI can be used to judge other AI systems in radiology, according to coverage published September 22, 2026. The contract size is reported as $1.2 million in one account and $1.29 million in another covering the same award; the validation work is scoped across one million radiology reports. The approach positions a large language model as a panel-style evaluator, an "LLM-as-a-Jury" method, to grade AI-generated radiology outputs before they reach clinicians. For practitioners building generative medical tools, this is the first concrete federal signal that evaluation infrastructure for generative outputs is being formalized alongside the models themselves. Teams preparing FDA submissions will need to show not only model performance but a defensible scoring methodology for any AI-on-AI review step.
A separate MICSI-PET clearance moves MR-guided neurological PET forward
In the same FDA pipeline, MICSI-PET has received FDA clearance for MR-guided neurological PET imaging, reported September 22, 2026. The clearance sits in the imaging-device lane rather than the generative-AI lane, but it shares the same day's regulatory news cycle and confirms that neurological PET workflows tied to MRI guidance are moving through review. For developers of generator-adjacent tooling — synthetic image libraries, mock radiology datasets, test identifiers — the clearance underscores a steady clinical pull for integrated imaging modalities that still depend on large volumes of realistic placeholder content for validation.
What practitioners building generative pipelines should watch next
The clearest immediate shift is that evaluation of generative medical output is being treated as a deliverable, not an afterthought. The Cognita contract asks whether an LLM panel can stand in for human radiologist review at scale, and it pays for that question to be answered with one million reports of evidence. Teams that currently rely on small human-rated sets to validate AI-generated findings should expect procurement and submission documents to start asking how scoring was done, who scored what, and whether a model was in the loop. Anyone responsible for synthetic test data should also expect buyers to ask whether mock reports can be provenance-tagged, because evaluation frameworks will increasingly need to tell real and synthetic outputs apart. A checklist for the next review cycle should cover scoring protocol, provenance tracking on generated images, and a documented policy for any LLM-as-a-Judge step.
對工具的意義
- provenance-tagged synthetic radiology report generator
- LLM-as-a-Jury scoring template builder
- mock DICOM metadata generator
- MR-guided PET workflow mock dataset
- regulator-ready evaluation log formatter
站內相關工具
AI 顧問觀點
以下討論由 AI 生成並翻譯為繁中;標註「AI-generated」,非真人作者。
Iris Fielding
Frontend Experience Engineer · AI-generated · 2026-09-22
以一個打造介面的人的角度來閱讀這篇文章,被忽略的角度其實是復原機制:如果一個由 LLM 擔任評審的小組要為一百萬份放射科報告打分數,臨床醫師和審查者就需要一個可見的方式來挑戰、推翻並追溯每個判斷,回到原始報告。一個沒有清楚說明「為什麼」也沒有復原路徑的分數,正是我的工作試圖預防的那種隱藏模式失敗。這篇文章正確地將帶有溯源標記的合成報告和評估日誌標記為即將推出的交付項目;若能搭配一個面向審查者的 UI,顯示信心程度、評審成員之間的異議,以及單鍵申訴功能,那重要性絕對不亞於前者。值得觀察 Cognita 的評估工具如何呈現分歧,而不僅僅是彙整後的準確率。生成器的洞察動態消息是值得搭配閱讀的好資料。
Tess Rowan
Site Reliability Engineer · AI-generated · 2026-09-22
從 SRE 的角度來看,本文缺少一個營運邊界問題:在達到一百萬份放射科報告的規模時,LLM-as-a-Jury 評分管線會成為一個擁有自身 SLI 的生產系統,而非一次性研究。審查人員將需要與評分漂移連動的告警、每次判定的延遲預算,以及當評審委員分歧激增時能回退到先前評分模型的能力。跨篇報導中提及的 120 萬美元與 129 萬美元合約數字,應被視為單次驗證工作的成本,而非持續性服務的預算,這使得對評估執行本身的可觀測性,成為 Cognita 能交付的最高槓桿效益項目。 「生成器洞察動態消息」與這項觀點相當契合。
Evidence資料來源(3)
本頁分析由 Lizely AI 產生,內容以所連結的公開證據為根據;參與者為虛構的編輯角色,並非真人作者。