Draft:HippoML
Draft article not currently submitted for review.
This is a draft Articles for creation (AfC) submission. It is not currently pending review. While there are no deadlines, abandoned drafts may be deleted after six months. To edit or make changes to this draft, simply click on the "Edit" tab at the top of the window. To be accepted, a draft should:
It is strongly discouraged to write about either yourself or your business or employer. If you do so, you must declare it. Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
Last edited by OpalYosutebito (talk | contribs) 3 months ago. (Update) |
| Type | Private (acquired) |
|---|---|
| Industry | Artificial intelligence, Machine learning, Computer software |
| Founded | 2023 |
| Founders | Bing Xu, Hao Lu, Terry Chen |
| Defunct | 2024 |
| Headquarters | Bellevue, Washington, United States |
Key people | Bing Xu (CEO) Terry Chen (VP of Engineering) Hao Lu |
| Products | HippoEngine, PrivateCanvas |
| Parent | NVIDIA |
HippoML, Inc. was an American artificial intelligence software company that developed inference engines and optimization tools for generative AI models.[1] Founded in 2023 and based in Bellevue, Washington, the company focused on improving deployment efficiency for large language models and other AI workloads across NVIDIA, AMD, and Apple Silicon hardware.[1][2][3] In 2024, HippoML was acquired by NVIDIA.[4]
History
HippoML was founded in Bellevue, Washington, in 2023 by Bing Xu, Hao Lu, and Terry Chen.[1][5][6][7] Before founding the company, Xu and Lu had worked on deep learning systems including AITemplate, while Chen had also contributed to GPU optimization frameworks including AITemplate.[5][6][7]
During its independent operation, HippoML published engineering results on model serving, attention optimization, quantized inference, and local AI applications.[2][8][9][10][3] In 2024, the company was acquired by NVIDIA.[4]
Technology and products
HippoML's primary technology was HippoEngine, a GPU inference engine designed to compile machine learning models ahead of time into standalone binary code rather than depend on large runtime environments. HippoML said the engine used a unified runtime and API with support for NVIDIA CUDA, AMD ROCm, and Apple Metal.[2]
HippoML also described a multiple layers model caching system for managing model weights across NVMe SSDs, host memory, and GPU memory. According to the company, this reduced model activation times to about 450 milliseconds.[2]
The company published several blog posts about transformer attention optimization, including variable-length attention, Apple Silicon acceleration, 8-bit HippoAttention, and claims of achieving 1 petaFLOPS attention performance on a single NVIDIA H100 SXM GPU.[8][9][10]
To demonstrate its inference stack on consumer hardware, HippoML released PrivateCanvas, a local desktop AI application for Windows, Linux, and macOS. HippoML described it as an offline tool for image generation and editing that bundled models such as Stable Diffusion XL, Segment Anything, and other generative models into one environment.[3]
HippoML also wrote about decentralized AI inference and the use of local devices for AI workloads in a February 2024 blog post.[11]
Post-acquisition
Following the acquisition, former HippoML founders and engineers appeared as authors on NVIDIA Technical Blog posts related to AI inference, including posts on automated GPU kernel generation and DeepSeek-R1 inference performance in 2025.[5][6][7][12][13]
See also
References
- ^ a b c "HippoML company profile". PitchBook. Retrieved March 30, 2026.
- ^ a b c d "Unified DataCenter & Local Foundation Model Serving: Beyond Docker Way". HippoML Blog. January 8, 2024. Retrieved March 30, 2026.
- ^ a b c "Super AI Creativity App Run with Local GPU on Mac/Windows/Linux [Early Access]". HippoML Blog. January 2, 2024. Retrieved March 30, 2026.
- ^ a b "Zheng Zhou". Wilson Sonsini. Retrieved March 30, 2026.
- ^ a b c "Bing Xu". NVIDIA Technical Blog. Retrieved March 30, 2026.
- ^ a b c "Hao Lu". NVIDIA Technical Blog. Retrieved March 30, 2026.
- ^ a b c "Terry Chen". NVIDIA Technical Blog. Retrieved March 30, 2026.
- ^ a b "Up to 80X Speedup in Multi-Head Attention on Apple Silicon". HippoML Blog. December 11, 2023. Retrieved March 30, 2026.
- ^ a b "8bit HippoAttention: Up to 3X Faster Compared to FlashAttentionV2". HippoML Blog. January 17, 2024. Retrieved March 30, 2026.
- ^ a b "PetaFLOPS Inference Era: 1 PFLOPS Attention, and Preliminary End-to-End Results". HippoML Blog. February 7, 2024. Retrieved March 30, 2026.
- ^ "Vision Pro, Decentralized GenAI and AI PC". HippoML Blog. February 7, 2024. Retrieved March 30, 2026.
- ^ "Automating GPU Kernel Generation with DeepSeek-R1 and Inference Time Scaling". NVIDIA Technical Blog. February 12, 2025. Retrieved March 30, 2026.
- ^ "NVIDIA Blackwell Delivers World-Record DeepSeek-R1 Inference Performance". NVIDIA Technical Blog. March 18, 2025. Retrieved March 30, 2026.
Content Disclaimer
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.
- The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
- There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
- It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
- Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
- Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.
