Open weights still have weight
1. 서로 다른 두 경로
AI 모델은 제공업체 서버에서 API로 호출하거나, Open Weight 모델을 내려받아 직접 실행한다. 중국의 AI 육성 전략과 미국의 수출통제가 이 제공 방식, 나아가 메모리 수요까지 바꿀까. 정책에서 메모리까지 이어지는 고리 중 어디까지가 관측이고 어디서부터가 추정인지 살펴본다.
- 미국 API 제공·중국 가중치 공개 : 정치적 개방성 아닌 산출물·제공 방식의 기술적 차이
- 양국 반례를 통해 관측 단위는 국가 성격 아닌 기업별 모델·제공 방식
2. 가중치에서 만나는 실행 통제
API로 모델을 쓸 때 어느 버전을 어디에서 어떤 소프트웨어로 돌릴지는 사용자가 아니라 모델 제공업체가 정한다. 오픈 웨이트 모델을 내려받으면 이 선택권도 함께 따라온다고 보기 쉽지만, 다운로드만으로 실제 운영 환경이 생기지는 않는다. 오픈 웨이트 모델을 자체 인프라나 선택한 호스팅 환경에 실제로 배포하면, 최종 사용자 또는 호스팅 제공업체가 모델 버전과 구동 위치, 런타임, 데이터 경로를 결정하고 인프라 운영 책임을 진다.
- 체크포인트 확보를 기반으로 자체·타사 호스팅 환경에 실제 배포 시 실행 환경 결정345
- 체크포인트 확보로 모델 수정·양자화 및 런타임 선택 가능
- 오픈 웨이트 모델 다운로드만으로 실제 배포·지속 사용 성립 불가
- 자체·타사 호스팅 환경에 실제 배포 시 최종 사용자 또는 호스팅 제공업체가 모델 버전·구동 위치·런타임·데이터 경로 결정
- 최종 사용자 또는 호스팅 제공업체 : 실행 환경 결정권과 인프라 운영 책임
- 메모리 : 가중치·KV 캐시·런타임 작업 공간 수용
- 연산과 소프트웨어 : 프리필·디코딩·도구 호출·서빙 런타임 운영
- 운영 : 전력·냉각·보안·검증·장애 대응 부담
- 핵심 교환 : 모델 제공업체 의존 축소와 인프라 운영 책임 부담
- 2025년 AI 행동계획의 오픈 모델 가치 명시와 별개로 공개 완화 체계는 공식 정책 미확인9
- 2025년 AI 행동계획에서 상업·학술·지정학적 가치 명시와 미국 선도 오픈 모델 지원
- 공개 여부와 방식은 개발자 결정으로 유지
- 2026년 공개 완화 체계는 단일 보도에서 논의 단계로 언급, 공식 정책 미확인10
측정 기준에 따른 겉보기 선두 역전 : 합산으로 시장 채택 규모·지속 사용 확정 불가
CAISI는 출시 후 경과 시점이 대략 비슷한 조건에서 DeepSeek V3.1과 gpt-oss를 비교. DeepSeek는 Hugging Face 다운로드와 파생 모델 업로드에서 뒤졌으나 OpenRouter 호스팅 요청에서는 앞섬. 각 지표는 플랫폼·측정 단위·기간·집계 범위가 다르므로, 이를 합쳐 오픈 웨이트 모델이 시장에서 얼마나 널리 채택되고 지속적으로 사용되는지 판단할 수 없다.11
- 리더십이 프론티어·유통·생태계로 분리되며 한 지표 선두만으로 배포 규모 확정 불가
- 프론티어 리더십 : 고정된 평가에서 측정한 모델 역량
- 유통 리더십 : 다운로드·호스팅 요청·파생 모델 등 서로 다른 활동 지표
- 생태계 리더십 : 런타임·도구·통합·운영 경험을 통한 지속 사용 가능성
- 한 지표의 선두로 다른 리더십이나 실제 배포 규모 확정 불가
3. 오픈 웨이트에서 추론 인프라로
오픈 웨이트 모델을 자체 또는 타사 환경에서 운영하면 추론 서버와 메모리도 그만큼 더 필요해질 것이라고 생각하기 쉽다. 그러나 내려받은 모델이 실제 운영에서 추론 요청을 처리하는지, 그 부하로 설비를 새로 늘리는지, 늘어난 설비에 어떤 메모리가 실리는지는 각각 따로 관측해야 하는 별개의 문제다. 하나가 확인되어도 그다음이 자동으로 따라오지 않고, 뒤에서 확인된 것이 앞의 빈자리를 메워 주지도 않는다. 관측이 멈춘 지점을 건너뛰고 이어 붙인 결론은 관측이 아니라 추정이다.
- 오픈 웨이트 모델의 자체·타사 호스팅 가능성 확립 : 실제 배포·지속 사용·추가 조달 없이 서버 출하 증가 확정 불가5
- 앞 절의 확립 범위 : 자체·타사 호스팅 가능성, 실제 배포 시 실행 환경 결정권과 인프라 운영 책임
- 수요 전환 조건 : 실제 외부 배포·지속 사용, 기존 인프라 용량을 초과하는 추가 조달
- 해석의 한계 : 다운로드·자체 호스팅 가능성만으로 서버 출하 증가 확정 불가
| 전환 조건 | 현재 결과 | 추가로 필요한 근거 |
|---|---|---|
| 자체·타사 호스팅 가능성 | 확립 | 실제 외부 배포·지속 사용 관측 |
| 실제 외부 배포 | 공개 사례에서 확립·시장 범위 미확인 | 모델별 활성 배포 수 |
| 지속 사용 | 공개 사례에서 확립·절대량 미공개 | 모델별 트래픽·지속 가동률 |
| 추가 서버 용량 | 사업자 증설 보고·오픈 웨이트 기여 미확인 | 모델별 신규 랙·클라우드 인스턴스·서버 조달 |
| 시스템당 DDR5 탑재량 | 대용량 계층 제품화·증분 미확인 | LPX 출하 구성·동일 워크로드 기준 시스템 BOM |
| CXMT 제품 공급 | 미확인 | 고객사 인증·고객별 제품 믹스·출하량·ASP·마진 |
- 배포 결과에 따른 추론 서버 수요 세 갈래
- 추가 수요 : 신규 또는 API에서 이전된 워크로드로 목적지 서버·클라우드 인스턴스 증설 유발 가능
- 수요 재배치 : 모델 제공업체 API에서 기업·클라우드·국가 주도 AI 사업자로 실행 주체 이동, 다만 기존 API 제공업체의 인프라 축소를 포함한 순증 효과는 미확인
- 수요 흡수 : 기존 서버 여유분·클라우드 용량·효율화로 신규 하드웨어 조달 없이 워크로드 수용
- 공개 사례의 실제 외부 배포·지속 사용 확립 : 모델별 기여도 미공개로 서버 순증 미확인1213141516
- 판정 기준 : 배포 수가 아니라 지속 가동률·추가 인프라 용량 동시 확인
- 실제 외부 배포 : Delhivery는 외부 API 대신 Llama 기반 오토스케일링 EC2 G5 인스턴스에 배포
- 수요 재배치 사례 : Gumloop은 전사 업무 에이전트를 폐쇄형 API 모델에서 오픈 웨이트 모델로 이전
- 지속 사용 : Factory에서 6개월간 오픈 웨이트 사용 비중 2~3배 확대
- 증설 사례 : 오픈 모델 사업자 GPU 사용 시간·데이터센터 용량 확대, 모델별 기여도 미공개
- LPX가 제품화한 DDR5 대용량 계층의 동일 워크로드 탑재 증분·HBM 대체량 미확인191720
- NVIDIA Groq 3 LPX 완전 구성 랙 광고 사양 128 GB S램·최대 12 TB DDR5를 대규모 모델·워크로드용 용량 계층으로 명시17
- 연산 분담 : Rubin GPU는 프리필과 누적 KV 캐시 대상 전체 컨텍스트 어텐션, Groq LPU는 FFN·MoE21
- KV 배치 : Pod 단위 KV 캐시 저장을 별도 BlueField-4 STX 랙에 배정, LPX DDR5에 저장되는 텐서·접근 빈도 미공개20
- 시스템 관계 : LPX는 Rubin NVL72 단독 대체가 아닌 병행 배치, 고정 배치 비율 미공개로 랙당 최대 20.7 TB HBM4와 합산 불가2022
- 대조 근거 : 2026년 하반기 공급 예정으로 출하 구성·DDR5 공급사 미공개, 동일 워크로드 기준 시스템 BOM 미공개로 증분·HBM 대체량 미확인20
- 근거 범위 : 검토 자료에서 LPX 채택·오픈 웨이트 배포 직접 연결 근거 부재로 HBM 대비 수요 비중이 아닌 제품 방향에 한정
- 대역폭·용량 조건에 따른 범용 D램 적격성이 분기되며 판정에 런타임 측정 필요
- 배포 규모 가설·탑재량 제품 방향 수준·CXMT 점유율 미확인으로 수요 산정 불가5111718
- 배포 규모 : 운영 환경에서 실제 추론 요청을 처리하는 모델 배포 수와 추가 인프라 용량
- 탑재량 : 추론 배포당 DDR5 용량
- CXMT 제품 비중 : 오픈 웨이트 추론 인프라에 공급되는 DDR5 중 CXMT 제품 비중
- 산식 : 추론 배포에서 발생하는 CXMT DDR5 수요 = 추론 배포 수 × 배포당 DDR5 탑재량 × CXMT 제품 비중
- 배포 규모 가설 : 자체·타사 호스팅 가능성에 따른 실제 추론 운영 확대 가능성, 규모 미확인
- 탑재량 근거 : LPX가 제시한 DDR5 대용량 계층의 제품 방향
- CXMT 비중 판단 : 오픈 웨이트 추론 인프라 내 CXMT 채택과 공급 비중 미확인
- 산식 범위 : CXMT의 PC·모바일·비AI 서버 물량 제외로 산식 결과와 CXMT 전체 DDR5 공급량 불일치
4. 두 번째 경쟁
수요가 실제로 늘어나는지와 별개로, 모델을 고르는 일은 그 모델을 대신 돌려 줄 모델 제공업체를 고르는 일이기도 하다. 두 번째 경쟁은 앞의 경쟁을 끝내고 들어서는 것이 아니라 그 위에 하나 더 열리며, 여기서는 어느 API를 부를지가 아니라 모델을 무엇 위에서 돌릴지를 고른다. 그 선택을 놓고 겨루는 쪽도 모델을 만드는 회사를 넘어선다. 다만 누가 이기는지는 여기서 답할 수 있는 물음이 아니다.
- 두 번째 경쟁 단위 : 모델 단독 아닌 가중치·가속기·메모리·런타임·도구·운영 결합
- 성능·비용과 하드웨어·런타임 등 문턱 동시 충족 시 오픈 웨이트가 대체재로 성립
- 대체 경로 확대 : 오픈 웨이트만으로 비용 절감·가격 협상력·시장 채택 규모 확정 불가
- 채택 조건 충족·지속 사용 관측을 통한 반복 채택 판정과 중단 위험 이동
- 작업 적격성 : 목표 작업의 성능·비용·지연 시간·신뢰성 충족
- 이전 가능 여부 : 모델·런타임·하드웨어의 호환성과 전환 부담
- 운영 가능 여부 : 개발 도구·문서·지원과 가속기·메모리 조달의 지속성
- 반복 채택 판정 : 일회성 실행을 넘어 실제 운영에서 사용이 이어지고 이를 뒷받침하는 조달·운영·지원도 지속되는지 확인
- 중단 위험의 이동 : 모델 제공업체 API 변경·중단 노출은 축소, 가속기·메모리·전력·운영 장애 노출은 사용자 측으로 이전
- 수출통제의 특정 첨단 연산·인터커넥트 하드웨어 접근 제약 : 대체 스택 성공의 증거 아님
- 중국 내 실행 스택 구성·범용 D램 공급 관측 : 성능·적격성은 측정·비교 부재로 판정 보류
- Huawei Ascend·CANN : 하드웨어와 드라이버·런타임·연산자·API를 잇는 구성 관측3435
- Huawei 판정 보류 : 공개 자료에서 성능·가용성·개발 경험 비교 부재34
- CXMT 범용 D램 : DDR4·DDR5·LPDDR4X·LPDDR5/5X 공급에 따른 메모리 선택지 확대1836
- CXMT 적격성 판정 보류 : 목표 작업별 런타임 측정과 수율·판매 가능 비트 관측값 부재3625
- 고성능 메모리 제약 : 2026년 7월 18일 공개 카탈로그에 HBM 제품 부재, HBM3 양산은 계획 단계로만 보도1837
- 양쪽의 이동 : 미국 오픈 모델 축 강화, 중국 가중치·국내 실행 기반 병행으로 전략적 선택지 확대95343418
- 결론의 범위 : 두 번째 경쟁의 성립과 판정 기준, 승자·시장 규모·메모리 업체별 수혜는 미확정
두 번째 경쟁에서 중요한 것은 프론티어 모델을 구동하는 AI 스택 전체를 그대로 재현하는 것이 아니라, 필요한 성능을 내는 대체 스택이 실제 운영에서 거듭 선택되는 것이다.
Sources and supplementary boundaries
- OpenAI, OpenAI API, June 2020. OpenAI described an API product intended to fund continued work, make expensive models accessible, and retain the ability to respond to misuse. Source
- OpenAI, GPT-4 Technical Report, March 2023. The report withheld architecture, hardware, training-compute, and dataset-construction details while providing model access through products and APIs. Source
- DeepSeek, DeepSeek-R1 repository. The repository links downloadable model weights and licenses the code and the R1 model weights under the MIT License, permitting commercial use, modification, and derivative works; the distilled variants carry their base models' licenses (Apache 2.0 for the Qwen-derived models, the Llama 3.1 and 3.3 licenses for the Llama-derived models). Source
- Alibaba Qwen team, Qwen3 repository. The official repository publishes downloadable model weights and documents local deployment paths. Source
- OpenAI, Introducing gpt-oss. OpenAI released open-weight models designed to run locally, on-device, or through third-party inference providers. Source
- Alibaba Cloud, Model Studio overview, last updated 2026-07-10. Model Studio serves the proprietary Qwen series and third-party models including DeepSeek and Kimi through hosted APIs billed only on invocation; downloadable Qwen weights are documented separately in the Qwen3 repository. Source
- State Council of China, New Generation Artificial Intelligence Development Plan, July 2017. The plan predates the current generative-AI cycle and the 2022 US export controls; it calls for stronger domestic AI capability and industrial development alongside open, collaborative innovation and shared foundational technologies. Source
- US Department of Commerce, Bureau of Industry and Security, October 2022 export controls on advanced computing and semiconductor manufacturing items to China. Source
- White House, America's AI Action Plan, July 2025. Its open-model section names commercial, academic, and geostrategic value and leaves release to the developer. Source
- Benjamin Guggenheim, AI & Tech Brief: Exclusive | An open-source framework, Washington Post Intelligence, July 13, 2026. The Post reports that the administration and the AI industry have been discussing a capability framework for US open-source models keyed to the capabilities of leading Chinese open-source models; no public order or framework confirms the mechanism. Source
- NIST Center for AI Standards and Innovation, 2025 evaluation of DeepSeek models. The report does not publish the underlying snapshot dates or complete raw table. Source
- AWS Delhivery case study, accessed 2026-07-29. AWS reports that Delhivery evaluated third-party serverless LLM APIs, rejected them on 2,000-requests-per-minute rate caps and provisioned-access cost, and instead deployed a fine-tuned open-source Llama 3.2 1B model to production on Amazon EKS with autoscaled EC2 G5 nodes, reaching up to 8,000 requests per minute at 160 ms. Source
- Fireworks AI Gumloop case study. Gumloop reports moving an internal production agent from Claude Opus to GLM and a sevenfold increase in open-weight agent chats over three weeks. Source
- Fireworks AI Factory case study. Factory reports that the open-weight share of its model usage increased two to three times over six months; absolute traffic and hardware capacity are not disclosed. Source
- AWS Simplismart case study. AWS reports an eightfold increase in deployed GPU-hours over three months, but the disclosed pool combines open, custom, multimodal, fine-tuning and inference workloads. Source
- Together AI, NVIDIA Cloud Partner announcement, March 2025. Together reports tens of thousands of deployed NVIDIA data-center GPUs and more than 200 MW of data-center and power capacity; the disclosure covers training and inference together and does not state the inference share, model-specific capacity, utilization, or retired capacity. Source
- NVIDIA, Groq 3 LPX product page. NVIDIA advertises 128 GB of SRAM and 12 TB of DDR5 per fully configured rack; in NVIDIA's separate tray specification the two DRAM paths are each qualified as up to. Source
- CXMT product catalog, checked 2026-07-18. The public catalog listed DDR4, DDR5, LPDDR4X and LPDDR5/5X but no HBM product. Source
- Groq, December 2025. Groq and NVIDIA entered a non-exclusive inference-technology licensing agreement; Groq stated that it would continue as an independent company. Source
- NVIDIA, Vera Rubin platform announcement, March 2026. NVIDIA introduced Groq 3 LPX as a rack-scale inference accelerator deployed with Vera Rubin NVL72 rather than as a standalone replacement for Rubin GPUs, said LPX racks would be available in the second half of 2026, and markets the separate BlueField-4 STX storage rack as the shared tier for storing and retrieving large KV-cache data across a pod. Source
- NVIDIA technical blog, Inside NVIDIA Groq 3 LPX. Rubin GPUs retain prefill and, during decode, full-context attention over the accumulated KV cache, while LPX executes latency-sensitive FFN and MoE operations in a heterogeneous inference system. Source
- NVIDIA Vera Rubin NVL72 preliminary specifications. NVIDIA lists up to 20.7 TB of HBM4 and 54 TB of LPDDR5X CPU memory per rack; values are preliminary and subject to change. Source
- Autoregressive decoding is often memory-bandwidth-bound because each token reads active weights at low arithmetic intensity. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are all you need, arXiv 2507.14397v1. Source
- NVIDIA TensorRT-LLM KV-cache documentation. Cache capacity and cross-request reuse depend on datatype, memory allocation, attention-window size, block reuse and eviction policy, scheduling, and runtime configuration. Source
- Memory-Bound but Not Bandwidth-Limited, arXiv 2605.30571. The fetched preprint finds memory-dominated batch-1 decode but shows that kernel-launch and runtime overhead can prevent realized performance from scaling with peak bandwidth. Source
- NVIDIA H200 product page specifications: 141 GB of HBM3e at 4.8 TB/s. Source
- DDR5 SDRAM. Each module presents a 64-bit channel split into two independent 32-bit sub-channels, and the specification table gives per-module bandwidth of 32.0-70.4 GB/s, i.e. data rate x 8 bytes. Two DDR5-6000 modules therefore give 2 x 6000 MT/s x 8 B = 96 GB/s; that product is this Leaf's arithmetic, not a figure the source states. Nominal peak, not measured application throughput. Source
- DeepSeek, V3 technical report: 671 billion total parameters, 37 billion activated per token. Source
- US National Telecommunications and Information Administration, Dual-Use Foundation Models with Widely Available Model Weights, July 2024, pages 8-9 and 29-33. The directional policy report discusses downstream customization, local or cloud execution, competition, and innovation; it does not establish eliminated provider dependence, universal cost savings, or operational parity. Source
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023, proceedings pages 611-626, sections 3-4 and 6. On identical hardware, PagedAttention reduces KV-cache fragmentation and vLLM reports 2-4x higher serving throughput than the evaluated systems; the paper does not establish escape from physical memory bandwidth or capacity limits. Source
- OpenAI, open-weight model documentation. Self-hosting makes the deployer responsible for compute, storage, and third-party hosting costs. Source
- Bresnahan and Greenstein, Technological Competition and the Structure of the Computer Industry, The Journal of Industrial Economics 47(1), 1999, pages 1-40. The historical analysis covers vertical disintegration, compatibility, horizontal competition, and shifts of control between computer-industry layers; it is an analogy, not evidence about modern AI stacks. Source
- West, How open is open enough? Melding proprietary and open source platform strategies, Research Policy 32(7), 2003, pages 1259-1285. The historical analysis covers Unix, open systems, and hybrid openness strategies; it does not establish portability across tightly coupled AI hardware and software. Source
- Huawei 2025 Annual Report. Huawei reports expanding Ascend-based AI infrastructure; it is a company disclosure, not an independent parity benchmark. Source
- Huawei CANN documentation, checked 2026-07-29. Huawei lists CANN's core components as driver, runtime, operator libraries, communication library, graph engine and compiler, with AscendCL as the C API layer for runtime management, single-operator invocation and model management. Source
- CXMT, Shanghai Stock Exchange STAR Market prospectus (registration draft), 27 May 2026. It names major end customers including Alibaba Cloud, ByteDance, Tencent and Lenovo, reached mainly through distributors; describes DDR and LPDDR series including server RDIMM and MRDIMM modules and proceeds earmarked for DRAM capacity and technology projects; and reports that DRAM revenue from AI compute servers remained a low share of the reporting period. It does not disclose customer-level product mix, absolute wafer capacity, a yield series, or saleable-bit output - the matched series an incumbent comparison would require. Source
- Tom's Hardware, on CXMT's reported plan to begin domestic HBM3 mass production by the end of 2026. This is reported planning, not a shipping-product disclosure. Source