為何愈來愈多企業希望在內部運行AI?

愈來愈多企業希望在內部運行AI(圖/達志影像美聯社)
對資料掌控、成本與速度的考量,讓在企業自有硬體上運行的AI系統備受關注。
一家製造商希望打造一個AI助理,能根據數十年累積的設計文件、維修紀錄與供應商合約,回答工程師的問題。若將這些檔案傳送至外部雲端服務,便會引發機密性、成本,以及資訊最終由誰掌控等疑慮。在公司自有伺服器上運行模型,則提供了另一條路。這樣的選擇,正是地端AI(on-premises AI)日益受到關注的原因。
地端AI是指在組織自行擁有並營運的硬體上運行AI模型,地點可能是自家資料中心、工廠或辦公室。公有雲服務仍是最容易入門的方式,因為企業可以租用運算資源,並依使用量付費。然而,隨著AI從實驗階段走入日常營運,愈來愈多組織開始重新思考這些工作應該在哪裡進行。
調查數據顯示市場興趣明顯升溫,但並非單向轉移。儲存業者Cloudian委託進行的調查,於今年2月訪問203位企業IT決策者,其中93%表示已將AI工作負載移出公有雲、正在移轉,或正在評估移轉。考量到委託業者本身經營地端儲存業務,這項數據應審慎看待。多數企業可能會同時使用雲端服務與自有基礎設施。
資料掌控是最常被提及的效益。同一項調查中,91%的受訪者表示,當AI需要使用敏感資料時,會選擇地端、私有或混合式基礎設施。對銀行、醫院、律師事務所與製造商而言,將客戶紀錄或專有設計保留在自家系統內,可簡化法規遵循,並降低資料暴露於外部單位的風險。美國參議院聯邦信用合作社(US Senate Federal Credit Union)正在打造首個用於資安的AI代理,基於安全與掌控考量,選擇在地端運行。
成本是另一項動機。雲端AI服務通常依使用量計費,以token計算,也就是模型處理文字的單位。這種模式適合實驗,但當員工與自動化代理全天使用AI後,支出便更難預測。在Cloudian的調查中,40%的受訪者表示雲端AI支出超出預期。自購硬體需要龐大的前期投資,但對穩定且大量的工作負載而言,成本可望更容易掌握。
速度同樣重要,尤其是在工廠。四分之三的受訪者表示,部分工作負載,例如即時影像分析與製造品質檢測,需要地端基礎設施才能達到足夠的反應速度。規模較小的專用模型,也讓在地部署更加可行。美國啤酒商新比利時釀酒公司(New Belgium Brewing)正以自家資料訓練模型,將部分釀造流程自動化,並認為單一前沿模型並不適用於所有任務。
相關硬體如今已從桌面延伸至資料中心。華碩的Ascent GX10與技嘉的AI TOP ATOM皆採用輝達精巧的GB10晶片,讓開發人員能在桌上運行大型模型。華碩的ExpertCenter Pro ET900N G3則採用輝達GB300晶片,搭配748GB記憶體,供企業在地端開發AI應用。在企業資料中心方面,華碩、技嘉、英業達、和碩、雲達、緯創與緯穎皆提供輝達RTX PRO伺服器。台積電與鴻海都是早期採用者,鴻海將其用於模擬生產線,以及開發自主移動機器人。
地端AI也有其成本。企業必須在記憶體價格上漲之際採購硬體,提供足夠的電力與散熱,並聘用人員維護系統與更新模型。將資料保留在內部,也無法消除資安風險。內部AI代理仍需要受限的權限、操作紀錄,以及防範隱藏在文件中惡意指令的保護機制。
這不太可能是非此即彼的選擇。許多企業將繼續使用雲端服務處理最大型的模型與波動的需求,同時在自有硬體上運行敏感或持續性的工作負載。對台灣的伺服器與PC業者而言,這將開創大型雲端業者以外的市場。地端AI的效益,包括資料掌控、更可預測的成本與更快的反應速度,唯有在系統證明可靠、安全且易於管理時,才能說服企業投資。
Why More Companies Want to Run AI In-House
Concerns about data control, cost, and speed are drawing attention to AI systems that run on a company's own hardware.
A manufacturer wants an AI assistant that can answer engineers' questions using decades of design documents, maintenance records, and supplier contracts. Sending those files to an outside cloud service raises questions about confidentiality, cost, and who ultimately controls the information. Running the model on the company's own servers offers another route. That choice lies behind the growing interest in on-premises AI.
On-premises AI means running AI models on hardware that an organization owns and operates, whether in its own data center, a factory, or an office. Public cloud services remain the easiest way to start, because companies can rent computing capacity and pay according to use. As AI moves from experiments into daily operations, however, more organizations are reconsidering where that work should take place.
Survey data points to significant interest, though not a one-way shift. In a February survey of 203 enterprise IT decision-makers commissioned by storage company Cloudian, 93% said they had moved AI workloads away from public cloud, were doing so, or were evaluating the move. Given the sponsor's interest in on-premises storage, the figure deserves some caution. Most companies are likely to use a combination of cloud services and their own infrastructure.
Data control is the most frequently cited benefit. In the same survey, 91% said they would choose on-premises, private, or hybrid infrastructure when AI uses sensitive data. For banks, hospitals, law firms, and manufacturers, keeping customer records or proprietary designs within their own systems simplifies compliance and reduces exposure to outside parties. The US Senate Federal Credit Union, which is building its first AI agent for cybersecurity, chose to run it on-premises for security and control.
Cost is another motivation. Cloud AI services typically charge by usage, measured in tokens, the units of text a model processes. That suits experimentation, but spending becomes harder to predict once employees and automated agents use AI throughout the day. In the Cloudian survey, 40% said cloud AI spending had exceeded their projections. Owning hardware requires a large upfront investment, but it can make costs more predictable for steady, heavy workloads.
Speed matters, particularly in factories. Three-quarters of respondents said some workloads, such as real-time video analytics and manufacturing quality control, need on-premises infrastructure to respond quickly enough. Smaller, specialized models also make local deployment more practical. New Belgium Brewing, a US brewer, is developing models trained on its own data to automate parts of its brewing process, arguing that a single frontier model is not suited to every task.
The hardware now extends from desks to data centers. Asus's Ascent GX10 and Gigabyte's AI TOP ATOM, both built on Nvidia's compact GB10 chip, allow developers to run large models on a desk. Asus's ExpertCenter Pro ET900N G3 uses Nvidia's GB300 chip with 748GB of memory to develop AI applications on-premises. For corporate data centers, Nvidia's RTX PRO servers are available from Asus, Gigabyte, Inventec, Pegatron, QCT, Wistron, and Wiwynn. TSMC and Foxconn are among the early users, with Foxconn applying them to simulate production lines and develop autonomous mobile robots.
On-premises AI carries its own costs. Companies must buy hardware while memory prices are rising, supply adequate power and cooling, and employ staff to maintain systems and update models. Keeping data in-house does not remove security risks either. An internal AI agent still needs restricted permissions, records of its actions, and protection against hostile instructions hidden in documents.
The choice is unlikely to be all or nothing. Many companies will continue using cloud services for the largest models and fluctuating demand, while running sensitive or constant workloads on their own hardware. For Taiwan's server and PC makers, that creates a market beyond the largest cloud operators. The benefits of on-premises AI, including control over data, more predictable costs, and faster responses, will persuade companies to invest only if the systems prove reliable, secure, and manageable.









