binderys
Leaf 4

Analytical Leaf ·

Open weights still have weight

1. 서로 다른 두 경로

AI 모델은 제공업체 서버에서 API로 호출하거나, Open Weight 모델을 내려받아 직접 실행한다. 중국의 AI 육성 전략과 미국의 수출통제가 이 제공 방식, 나아가 메모리 수요까지 바꿀까. 정책에서 메모리까지 이어지는 고리 중 어디까지가 관측이고 어디서부터가 추정인지 살펴본다.

  • 미국 API 제공·중국 가중치 공개 : 정치적 개방성 아닌 산출물·제공 방식의 기술적 차이
    • 미국 경로의 한 축 : 대규모 훈련 역량과 모델 제공업체가 운영하는 API·클라우드 서비스의 결합12
    • 중국 경로의 한 축 : 다운로드·수정·재배포 가능한 DeepSeek-R1·Qwen3 가중치 공개34
    • 차이의 성격 : 정치적 개방성이 아니라 모델 산출물·제공 방식의 기술적 차이
  • 양국 반례를 통해 관측 단위는 국가 성격 아닌 기업별 모델·제공 방식
    • 미국 반례 : OpenAI의 gpt-oss 오픈 웨이트 공개와 자체 호스팅 지원5
    • 중국 반례 : Alibaba의 Qwen 오픈 웨이트와 독점형·종량제 Model Studio API 병행46
    • 관측 단위 : 국가 성격이 아니라 기업별 모델·라이선스·제공 방식
  • 중국 AI 육성 전략·미국 수출통제 : 두 경로와의 직접 인과 미확립
    • 중국은 2017년 ‘차세대 인공지능 발전계획’에서 국내 AI 역량 강화와 개방·공유를 함께 정책 목표로 제시7
    • 미국 첨단 연산 반도체 수출통제로 중국 기업의 가속기 접근 범위 축소8
    • 해석의 한계 : 자본·연산 접근성이 폐쇄형 제공을, 제약이 효율화·가중치 공개를 직접 유발했다는 인과는 미확립
    • 확립 가능한 출발점 : 중앙집중 제공과 가중치 유통 사례의 반복 관측

2. 가중치에서 만나는 실행 통제

API로 모델을 쓸 때 어느 버전을 어디에서 어떤 소프트웨어로 돌릴지는 사용자가 아니라 모델 제공업체가 정한다. 오픈 웨이트 모델을 내려받으면 이 선택권도 함께 따라온다고 보기 쉽지만, 다운로드만으로 실제 운영 환경이 생기지는 않는다. 오픈 웨이트 모델을 자체 인프라나 선택한 호스팅 환경에 실제로 배포하면, 최종 사용자 또는 호스팅 제공업체가 모델 버전과 구동 위치, 런타임, 데이터 경로를 결정하고 인프라 운영 책임을 진다.

  • 체크포인트 확보를 기반으로 자체·타사 호스팅 환경에 실제 배포 시 실행 환경 결정345
    • 체크포인트 확보로 모델 수정·양자화 및 런타임 선택 가능
    • 오픈 웨이트 모델 다운로드만으로 실제 배포·지속 사용 성립 불가
    • 자체·타사 호스팅 환경에 실제 배포 시 최종 사용자 또는 호스팅 제공업체가 모델 버전·구동 위치·런타임·데이터 경로 결정
  • 최종 사용자 또는 호스팅 제공업체 : 실행 환경 결정권과 인프라 운영 책임
    • 메모리 : 가중치·KV 캐시·런타임 작업 공간 수용
    • 연산과 소프트웨어 : 프리필·디코딩·도구 호출·서빙 런타임 운영
    • 운영 : 전력·냉각·보안·검증·장애 대응 부담
    • 핵심 교환 : 모델 제공업체 의존 축소와 인프라 운영 책임 부담
  • 2025년 AI 행동계획의 오픈 모델 가치 명시와 별개로 공개 완화 체계는 공식 정책 미확인9
    • 2025년 AI 행동계획에서 상업·학술·지정학적 가치 명시와 미국 선도 오픈 모델 지원
    • 공개 여부와 방식은 개발자 결정으로 유지
    • 2026년 공개 완화 체계는 단일 보도에서 논의 단계로 언급, 공식 정책 미확인10

측정 기준에 따른 겉보기 선두 역전 : 합산으로 시장 채택 규모·지속 사용 확정 불가

CAISI는 출시 후 경과 시점이 대략 비슷한 조건에서 DeepSeek V3.1과 gpt-oss를 비교. DeepSeek는 Hugging Face 다운로드와 파생 모델 업로드에서 뒤졌으나 OpenRouter 호스팅 요청에서는 앞섬. 각 지표는 플랫폼·측정 단위·기간·집계 범위가 다르므로, 이를 합쳐 오픈 웨이트 모델이 시장에서 얼마나 널리 채택되고 지속적으로 사용되는지 판단할 수 없다.11

세 개의 개별 비교. DeepSeek V3.1은 Hugging Face 다운로드가 gpt-oss-20b의 약 2%, OpenRouter 요청은 gpt-oss 패밀리보다 약 25% 많음, Hugging Face 파생 업로드는 gpt-oss-20b의 12% 미만.
  • 리더십이 프론티어·유통·생태계로 분리되며 한 지표 선두만으로 배포 규모 확정 불가
    • 프론티어 리더십 : 고정된 평가에서 측정한 모델 역량
    • 유통 리더십 : 다운로드·호스팅 요청·파생 모델 등 서로 다른 활동 지표
    • 생태계 리더십 : 런타임·도구·통합·운영 경험을 통한 지속 사용 가능성
    • 한 지표의 선두로 다른 리더십이나 실제 배포 규모 확정 불가

3. 오픈 웨이트에서 추론 인프라로

오픈 웨이트 모델을 자체 또는 타사 환경에서 운영하면 추론 서버와 메모리도 그만큼 더 필요해질 것이라고 생각하기 쉽다. 그러나 내려받은 모델이 실제 운영에서 추론 요청을 처리하는지, 그 부하로 설비를 새로 늘리는지, 늘어난 설비에 어떤 메모리가 실리는지는 각각 따로 관측해야 하는 별개의 문제다. 하나가 확인되어도 그다음이 자동으로 따라오지 않고, 뒤에서 확인된 것이 앞의 빈자리를 메워 주지도 않는다. 관측이 멈춘 지점을 건너뛰고 이어 붙인 결론은 관측이 아니라 추정이다.

  • 오픈 웨이트 모델의 자체·타사 호스팅 가능성 확립 : 실제 배포·지속 사용·추가 조달 없이 서버 출하 증가 확정 불가5
    • 앞 절의 확립 범위 : 자체·타사 호스팅 가능성, 실제 배포 시 실행 환경 결정권과 인프라 운영 책임
    • 수요 전환 조건 : 실제 외부 배포·지속 사용, 기존 인프라 용량을 초과하는 추가 조달
    • 해석의 한계 : 다운로드·자체 호스팅 가능성만으로 서버 출하 증가 확정 불가
전환 조건 여섯 단계 중 첫 단계 실행 가능성 확립 이후 단계마다 미확인·미공개 항목 잔존51112131415161718
전환 조건현재 결과추가로 필요한 근거
자체·타사 호스팅 가능성확립실제 외부 배포·지속 사용 관측
실제 외부 배포공개 사례에서 확립·시장 범위 미확인모델별 활성 배포 수
지속 사용공개 사례에서 확립·절대량 미공개모델별 트래픽·지속 가동률
추가 서버 용량사업자 증설 보고·오픈 웨이트 기여 미확인모델별 신규 랙·클라우드 인스턴스·서버 조달
시스템당 DDR5 탑재량대용량 계층 제품화·증분 미확인LPX 출하 구성·동일 워크로드 기준 시스템 BOM
CXMT 제품 공급미확인고객사 인증·고객별 제품 믹스·출하량·ASP·마진
  • 배포 결과에 따른 추론 서버 수요 세 갈래
    • 추가 수요 : 신규 또는 API에서 이전된 워크로드로 목적지 서버·클라우드 인스턴스 증설 유발 가능
    • 수요 재배치 : 모델 제공업체 API에서 기업·클라우드·국가 주도 AI 사업자로 실행 주체 이동, 다만 기존 API 제공업체의 인프라 축소를 포함한 순증 효과는 미확인
    • 수요 흡수 : 기존 서버 여유분·클라우드 용량·효율화로 신규 하드웨어 조달 없이 워크로드 수용
  • 공개 사례의 실제 외부 배포·지속 사용 확립 : 모델별 기여도 미공개로 서버 순증 미확인1213141516
    • 판정 기준 : 배포 수가 아니라 지속 가동률·추가 인프라 용량 동시 확인
    • 실제 외부 배포 : Delhivery는 외부 API 대신 Llama 기반 오토스케일링 EC2 G5 인스턴스에 배포
    • 수요 재배치 사례 : Gumloop은 전사 업무 에이전트를 폐쇄형 API 모델에서 오픈 웨이트 모델로 이전
    • 지속 사용 : Factory에서 6개월간 오픈 웨이트 사용 비중 2~3배 확대
    • 증설 사례 : 오픈 모델 사업자 GPU 사용 시간·데이터센터 용량 확대, 모델별 기여도 미공개
  • LPX가 제품화한 DDR5 대용량 계층의 동일 워크로드 탑재 증분·HBM 대체량 미확인191720
    • NVIDIA Groq 3 LPX 완전 구성 랙 광고 사양 128 GB S램·최대 12 TB DDR5를 대규모 모델·워크로드용 용량 계층으로 명시17
    • 연산 분담 : Rubin GPU는 프리필과 누적 KV 캐시 대상 전체 컨텍스트 어텐션, Groq LPU는 FFN·MoE21
    • KV 배치 : Pod 단위 KV 캐시 저장을 별도 BlueField-4 STX 랙에 배정, LPX DDR5에 저장되는 텐서·접근 빈도 미공개20
    • 시스템 관계 : LPX는 Rubin NVL72 단독 대체가 아닌 병행 배치, 고정 배치 비율 미공개로 랙당 최대 20.7 TB HBM4와 합산 불가2022
    • 대조 근거 : 2026년 하반기 공급 예정으로 출하 구성·DDR5 공급사 미공개, 동일 워크로드 기준 시스템 BOM 미공개로 증분·HBM 대체량 미확인20
    • 근거 범위 : 검토 자료에서 LPX 채택·오픈 웨이트 배포 직접 연결 근거 부재로 HBM 대비 수요 비중이 아닌 제품 방향에 한정
  • 대역폭·용량 조건에 따른 범용 D램 적격성이 분기되며 판정에 런타임 측정 필요
    • 낮은 동시성 자기회귀 디코딩에서 활성 가중치와 KV 캐시 읽기로 메모리 지배 가능2324
    • 희소성·양자화로 토큰당 읽기량 감소 가능
    • 명목상 최대치와 실제 지속 대역폭의 차이로 런타임 측정 필요25
    • HBM 고대역폭과 범용 DDR5 대용량 사이에서 워크로드별 적격성 분기2627
  • DDR5 모듈 2개 구성 단일 스트림 산술 상한(한 자릿수 토큰/초) : 실제 처리량·배포 가능성 미확립2827
    • 전체 파라미터 671B·토큰당 활성 파라미터 37B, 4비트 단순 환산 기준 단계당 활성 가중치 약 18~19 GB
    • DDR5-6000 모듈 2개·총 128비트 폭 이론 대역폭 약 96 GB/s 기준 단일 스트림 낙관적 상한 약 5 토큰/초
    • 4비트 덴스 70B 기준 DDR5 단일 스트림 약 2~3 토큰/초로, 희소성 없는 구성에서는 범용 D램 적격성 미성립
    • 산술 예시 한계 : KV 캐시·런타임·프리필·종단 간 지연 시간 제외로 실제 처리량·배포 가능성 미확립2524
  • 배포 규모 가설·탑재량 제품 방향 수준·CXMT 점유율 미확인으로 수요 산정 불가5111718
    • 배포 규모 : 운영 환경에서 실제 추론 요청을 처리하는 모델 배포 수와 추가 인프라 용량
    • 탑재량 : 추론 배포당 DDR5 용량
    • CXMT 제품 비중 : 오픈 웨이트 추론 인프라에 공급되는 DDR5 중 CXMT 제품 비중
    • 산식 : 추론 배포에서 발생하는 CXMT DDR5 수요 = 추론 배포 수 × 배포당 DDR5 탑재량 × CXMT 제품 비중
    • 배포 규모 가설 : 자체·타사 호스팅 가능성에 따른 실제 추론 운영 확대 가능성, 규모 미확인
    • 탑재량 근거 : LPX가 제시한 DDR5 대용량 계층의 제품 방향
    • CXMT 비중 판단 : 오픈 웨이트 추론 인프라 내 CXMT 채택과 공급 비중 미확인
    • 산식 범위 : CXMT의 PC·모바일·비AI 서버 물량 제외로 산식 결과와 CXMT 전체 DDR5 공급량 불일치

4. 두 번째 경쟁

수요가 실제로 늘어나는지와 별개로, 모델을 고르는 일은 그 모델을 대신 돌려 줄 모델 제공업체를 고르는 일이기도 하다. 두 번째 경쟁은 앞의 경쟁을 끝내고 들어서는 것이 아니라 그 위에 하나 더 열리며, 여기서는 어느 API를 부를지가 아니라 모델을 무엇 위에서 돌릴지를 고른다. 그 선택을 놓고 겨루는 쪽도 모델을 만드는 회사를 넘어선다. 다만 누가 이기는지는 여기서 답할 수 있는 물음이 아니다.

  • 두 번째 경쟁 단위 : 모델 단독 아닌 가중치·가속기·메모리·런타임·도구·운영 결합
    • 첫 번째 경쟁 : 가장 강한 모델과 중앙집중 인프라 보유12
    • 두 번째 경쟁 : 최종 사용자가 모델 제공업체 API에 의존하지 않고 채택·수정·운영 가능한 모델과 실행 스택 제공345
    • 경쟁 단위 : 모델 단독이 아니라 가중치·가속기·메모리·런타임·도구·운영의 결합
  • 성능·비용과 하드웨어·런타임 등 문턱 동시 충족 시 오픈 웨이트가 대체재로 성립
    • 오픈 웨이트 : 최종 사용자 또는 호스팅 제공업체의 실행 환경 결정에 필요, 실제 배포·지속 사용에는 불충분34529
    • 성립 문턱 : 목표 작업의 성능·비용·신뢰성과 하드웨어·메모리·런타임·도구의 동시 충족2524
    • 런타임의 역할 : 평가된 동일 하드웨어 조건에서 KV 캐시 메모리 효율화와 서빙 처리량 2~4배 개선30
    • 물리적 한계 : 소프트웨어 최적화로 제거되지 않는 가속기 메모리 용량·대역폭 제약2627
  • 대체 경로 확대 : 오픈 웨이트만으로 비용 절감·가격 협상력·시장 채택 규모 확정 불가
    • 가중치 확보로 중앙집중 API와 자체 실행 사이의 대체 경로 확대29
    • 대체 경로의 효과 : 후속 개발자의 수정·현장 적용 범위 확대29
    • 경제적 효과 : 단일 모델 제공업체 의존과 전환 부담의 축소 가능성29
    • 자체 실행의 비용 구조 : 연산·저장·외부 서비스 비용은 최종 사용자 또는 선택한 호스트 부담31
    • 결론의 한계 : 오픈 웨이트만으로 비용 절감·가격 협상력·시장 채택 규모 확정 불가291131
  • 채택 조건 충족·지속 사용 관측을 통한 반복 채택 판정과 중단 위험 이동
    • 작업 적격성 : 목표 작업의 성능·비용·지연 시간·신뢰성 충족
    • 이전 가능 여부 : 모델·런타임·하드웨어의 호환성과 전환 부담
    • 운영 가능 여부 : 개발 도구·문서·지원과 가속기·메모리 조달의 지속성
    • 반복 채택 판정 : 일회성 실행을 넘어 실제 운영에서 사용이 이어지고 이를 뒷받침하는 조달·운영·지원도 지속되는지 확인
    • 중단 위험의 이동 : 모델 제공업체 API 변경·중단 노출은 축소, 가속기·메모리·전력·운영 장애 노출은 사용자 측으로 이전
  • 수출통제의 특정 첨단 연산·인터커넥트 하드웨어 접근 제약 : 대체 스택 성공의 증거 아님
    • 수출통제의 의미 : 특정 첨단 연산·인터커넥트 하드웨어 접근 제약, 대체 스택 성공의 증거는 아님8
    • ※ 역사적 유비 1 : 컴퓨터 산업의 호환 인터페이스는 계층별 경쟁·주도권 이동을 촉진한 선례, 다만 현대 AI 스택에 같은 결과를 보장하지 않음32
    • ※ 역사적 유비 2 : Linux·POSIX는 독점형 유닉스 의존을 낮춘 대체 경로 선례, 다만 AI의 강한 하드웨어·소프트웨어 결합까지 설명하지 못함33
  • 중국 내 실행 스택 구성·범용 D램 공급 관측 : 성능·적격성은 측정·비교 부재로 판정 보류
    • Huawei Ascend·CANN : 하드웨어와 드라이버·런타임·연산자·API를 잇는 구성 관측3435
    • Huawei 판정 보류 : 공개 자료에서 성능·가용성·개발 경험 비교 부재34
    • CXMT 범용 D램 : DDR4·DDR5·LPDDR4X·LPDDR5/5X 공급에 따른 메모리 선택지 확대1836
    • CXMT 적격성 판정 보류 : 목표 작업별 런타임 측정과 수율·판매 가능 비트 관측값 부재3625
    • 고성능 메모리 제약 : 2026년 7월 18일 공개 카탈로그에 HBM 제품 부재, HBM3 양산은 계획 단계로만 보도1837
    • 양쪽의 이동 : 미국 오픈 모델 축 강화, 중국 가중치·국내 실행 기반 병행으로 전략적 선택지 확대95343418
    • 결론의 범위 : 두 번째 경쟁의 성립과 판정 기준, 승자·시장 규모·메모리 업체별 수혜는 미확정

두 번째 경쟁에서 중요한 것은 프론티어 모델을 구동하는 AI 스택 전체를 그대로 재현하는 것이 아니라, 필요한 성능을 내는 대체 스택이 실제 운영에서 거듭 선택되는 것이다.

Sources and supplementary boundaries

  1. OpenAI, OpenAI API, June 2020. OpenAI described an API product intended to fund continued work, make expensive models accessible, and retain the ability to respond to misuse. Source
  2. OpenAI, GPT-4 Technical Report, March 2023. The report withheld architecture, hardware, training-compute, and dataset-construction details while providing model access through products and APIs. Source
  3. DeepSeek, DeepSeek-R1 repository. The repository links downloadable model weights and licenses the code and the R1 model weights under the MIT License, permitting commercial use, modification, and derivative works; the distilled variants carry their base models' licenses (Apache 2.0 for the Qwen-derived models, the Llama 3.1 and 3.3 licenses for the Llama-derived models). Source
  4. Alibaba Qwen team, Qwen3 repository. The official repository publishes downloadable model weights and documents local deployment paths. Source
  5. OpenAI, Introducing gpt-oss. OpenAI released open-weight models designed to run locally, on-device, or through third-party inference providers. Source
  6. Alibaba Cloud, Model Studio overview, last updated 2026-07-10. Model Studio serves the proprietary Qwen series and third-party models including DeepSeek and Kimi through hosted APIs billed only on invocation; downloadable Qwen weights are documented separately in the Qwen3 repository. Source
  7. State Council of China, New Generation Artificial Intelligence Development Plan, July 2017. The plan predates the current generative-AI cycle and the 2022 US export controls; it calls for stronger domestic AI capability and industrial development alongside open, collaborative innovation and shared foundational technologies. Source
  8. US Department of Commerce, Bureau of Industry and Security, October 2022 export controls on advanced computing and semiconductor manufacturing items to China. Source
  9. White House, America's AI Action Plan, July 2025. Its open-model section names commercial, academic, and geostrategic value and leaves release to the developer. Source
  10. Benjamin Guggenheim, AI & Tech Brief: Exclusive | An open-source framework, Washington Post Intelligence, July 13, 2026. The Post reports that the administration and the AI industry have been discussing a capability framework for US open-source models keyed to the capabilities of leading Chinese open-source models; no public order or framework confirms the mechanism. Source
  11. NIST Center for AI Standards and Innovation, 2025 evaluation of DeepSeek models. The report does not publish the underlying snapshot dates or complete raw table. Source
  12. AWS Delhivery case study, accessed 2026-07-29. AWS reports that Delhivery evaluated third-party serverless LLM APIs, rejected them on 2,000-requests-per-minute rate caps and provisioned-access cost, and instead deployed a fine-tuned open-source Llama 3.2 1B model to production on Amazon EKS with autoscaled EC2 G5 nodes, reaching up to 8,000 requests per minute at 160 ms. Source
  13. Fireworks AI Gumloop case study. Gumloop reports moving an internal production agent from Claude Opus to GLM and a sevenfold increase in open-weight agent chats over three weeks. Source
  14. Fireworks AI Factory case study. Factory reports that the open-weight share of its model usage increased two to three times over six months; absolute traffic and hardware capacity are not disclosed. Source
  15. AWS Simplismart case study. AWS reports an eightfold increase in deployed GPU-hours over three months, but the disclosed pool combines open, custom, multimodal, fine-tuning and inference workloads. Source
  16. Together AI, NVIDIA Cloud Partner announcement, March 2025. Together reports tens of thousands of deployed NVIDIA data-center GPUs and more than 200 MW of data-center and power capacity; the disclosure covers training and inference together and does not state the inference share, model-specific capacity, utilization, or retired capacity. Source
  17. NVIDIA, Groq 3 LPX product page. NVIDIA advertises 128 GB of SRAM and 12 TB of DDR5 per fully configured rack; in NVIDIA's separate tray specification the two DRAM paths are each qualified as up to. Source
  18. CXMT product catalog, checked 2026-07-18. The public catalog listed DDR4, DDR5, LPDDR4X and LPDDR5/5X but no HBM product. Source
  19. Groq, December 2025. Groq and NVIDIA entered a non-exclusive inference-technology licensing agreement; Groq stated that it would continue as an independent company. Source
  20. NVIDIA, Vera Rubin platform announcement, March 2026. NVIDIA introduced Groq 3 LPX as a rack-scale inference accelerator deployed with Vera Rubin NVL72 rather than as a standalone replacement for Rubin GPUs, said LPX racks would be available in the second half of 2026, and markets the separate BlueField-4 STX storage rack as the shared tier for storing and retrieving large KV-cache data across a pod. Source
  21. NVIDIA technical blog, Inside NVIDIA Groq 3 LPX. Rubin GPUs retain prefill and, during decode, full-context attention over the accumulated KV cache, while LPX executes latency-sensitive FFN and MoE operations in a heterogeneous inference system. Source
  22. NVIDIA Vera Rubin NVL72 preliminary specifications. NVIDIA lists up to 20.7 TB of HBM4 and 54 TB of LPDDR5X CPU memory per rack; values are preliminary and subject to change. Source
  23. Autoregressive decoding is often memory-bandwidth-bound because each token reads active weights at low arithmetic intensity. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are all you need, arXiv 2507.14397v1. Source
  24. NVIDIA TensorRT-LLM KV-cache documentation. Cache capacity and cross-request reuse depend on datatype, memory allocation, attention-window size, block reuse and eviction policy, scheduling, and runtime configuration. Source
  25. Memory-Bound but Not Bandwidth-Limited, arXiv 2605.30571. The fetched preprint finds memory-dominated batch-1 decode but shows that kernel-launch and runtime overhead can prevent realized performance from scaling with peak bandwidth. Source
  26. NVIDIA H200 product page specifications: 141 GB of HBM3e at 4.8 TB/s. Source
  27. DDR5 SDRAM. Each module presents a 64-bit channel split into two independent 32-bit sub-channels, and the specification table gives per-module bandwidth of 32.0-70.4 GB/s, i.e. data rate x 8 bytes. Two DDR5-6000 modules therefore give 2 x 6000 MT/s x 8 B = 96 GB/s; that product is this Leaf's arithmetic, not a figure the source states. Nominal peak, not measured application throughput. Source
  28. DeepSeek, V3 technical report: 671 billion total parameters, 37 billion activated per token. Source
  29. US National Telecommunications and Information Administration, Dual-Use Foundation Models with Widely Available Model Weights, July 2024, pages 8-9 and 29-33. The directional policy report discusses downstream customization, local or cloud execution, competition, and innovation; it does not establish eliminated provider dependence, universal cost savings, or operational parity. Source
  30. Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023, proceedings pages 611-626, sections 3-4 and 6. On identical hardware, PagedAttention reduces KV-cache fragmentation and vLLM reports 2-4x higher serving throughput than the evaluated systems; the paper does not establish escape from physical memory bandwidth or capacity limits. Source
  31. OpenAI, open-weight model documentation. Self-hosting makes the deployer responsible for compute, storage, and third-party hosting costs. Source
  32. Bresnahan and Greenstein, Technological Competition and the Structure of the Computer Industry, The Journal of Industrial Economics 47(1), 1999, pages 1-40. The historical analysis covers vertical disintegration, compatibility, horizontal competition, and shifts of control between computer-industry layers; it is an analogy, not evidence about modern AI stacks. Source
  33. West, How open is open enough? Melding proprietary and open source platform strategies, Research Policy 32(7), 2003, pages 1259-1285. The historical analysis covers Unix, open systems, and hybrid openness strategies; it does not establish portability across tightly coupled AI hardware and software. Source
  34. Huawei 2025 Annual Report. Huawei reports expanding Ascend-based AI infrastructure; it is a company disclosure, not an independent parity benchmark. Source
  35. Huawei CANN documentation, checked 2026-07-29. Huawei lists CANN's core components as driver, runtime, operator libraries, communication library, graph engine and compiler, with AscendCL as the C API layer for runtime management, single-operator invocation and model management. Source
  36. CXMT, Shanghai Stock Exchange STAR Market prospectus (registration draft), 27 May 2026. It names major end customers including Alibaba Cloud, ByteDance, Tencent and Lenovo, reached mainly through distributors; describes DDR and LPDDR series including server RDIMM and MRDIMM modules and proceeds earmarked for DRAM capacity and technology projects; and reports that DRAM revenue from AI compute servers remained a low share of the reporting period. It does not disclose customer-level product mix, absolute wafer capacity, a yield series, or saleable-bit output - the matched series an incumbent comparison would require. Source
  37. Tom's Hardware, on CXMT's reported plan to begin domestic HBM3 mass production by the end of 2026. This is reported planning, not a shipping-product disclosure. Source