Base Compute
Base Compute is an AI inference lab building the runtimes and infrastructure that make powerful AI run on-device. Its stated mission is to run AGI on-device, so that inference happens fast, privately, and at near-zero marginal cost on hardware people and organisations already own.
At a Glance
- Enterprises in regulated industries
- Financial services
- Healthcare providers
- Legal firms
- +4 more
AI Tools by Base Compute
(1)BaseRT
LLM Inference Runtime for Apple Silicon
Discussions
No discussions yet
Be the first to start a discussion about Base Compute
Latest News
Automated Research: GLM 5.2 speeds up its own inference
GLM 5.2: near-frontier on our Mac Studio with 512GB RAM, no cloud involved
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
New performance ceiling for LLMs on Apple's M5 Pro
Products & Services
A native Metal LLM inference runtime for Apple Silicon, installed with a single shell command. Reports prefill up to 6.4x faster than llama.cpp and 3.9x faster than MLX, and decode up to 1.33x faster than MLX, using hand-written Metal 4 tensor-core kernels for dense and mixture-of-experts operations and M5 Neural Accelerators. Distributed openly via GitHub.
The management layer for AI across an organisation's device fleet, covering model catalogue, model distribution, local and cloud routing, role-based access, policy enforcement, and monitoring.
Applications built on the stack for enterprise workflows, including assistants, scribes, chat, voice, agentic flows, document workflows, video analytics, deep research, and coding agents. Organisations can also bring their own UI.
A simulation tool that models what on-device AI could deliver for a given organisation's use case and device fleet.
Market Position
Base Compute positions BaseRT against the two incumbent local-inference runtimes, llama.cpp and Apple's MLX, competing on raw throughput on Apple Silicon rather than on breadth of hardware support. Its broader argument is that open models now trail the closed frontier by roughly three months while running free on hardware an organisation already owns, which reframes on-device inference as an alternative to metered cloud APIs rather than a fallback for offline use.
Key Competitors
Leadership
Founders
Fabian Waschkowski
Listed as an author on both BaseRT technical reports, affiliated with Base Compute in Melbourne, Australia. Public sources do not state a formal title.
Prabod Rathnayaka
Listed as an author on both BaseRT technical reports, affiliated with Base Compute in Melbourne, Australia. Public sources do not state a formal title.
Lukas Wesemann
Listed as an author on both BaseRT technical reports and shows Base Compute on his LinkedIn profile. Also associated with the Australian ML/AI community organisation MLAI. Public sources do not state a formal title at Base Compute.
Founding Story
Base Compute was founded in 2026 around the argument that capable open models now fit on hardware people and companies already own, which changes where inference should run. The team describes itself as a small senior group in Melbourne and Berlin working at the runtime and silicon level. It published its position in an April 2026 post, 'Why we're betting on AGI that runs on-device', then shipped its first product, the BaseRT inference runtime, on July 1, 2026.
Business Model
Pricing Tiers
The BaseRT runtime is publicly available and installable via a single shell command, with source on GitHub. No price is published on the public site.
Enterprise stack covering Base Apps, Base Control, and BaseRT for on-premise, hybrid, or air-gapped deployment. Engagement runs through contact and the On-device AI Profiler. Specific pricing is not retrievable from the public site.
Target Markets
- Enterprises in regulated industries
- Financial services
- Healthcare providers
- Legal firms
- Government and defence
- Engineering organisations
- Running local coding assistants on developer machines to cut monthly AI spend
- Meeting assistants, drafting, and analysis over market-sensitive financial information
- Clinical documentation and patient-record summarisation on a clinician's own device
- Legal research, drafting, and review across privileged documents that cannot leave the firm
- Transcription and analysis of sensitive government and defence material, up to fully air-gapped
- Real-time robotics perception and control without a network round-trip