MLX on a Mac
Apple-silicon model serving via MLX — CLI server or the oMLX menu-bar app · AI / LLMs
MLX runs large language models natively on Apple Silicon — unified M-series memory, often faster and lighter than llama.cpp on a Mac, with models from the mlx-community org on Hugging Face. FrontierStack manages both ways to serve them, in this one pane: MLX-LM (pip install mlx-lm, then mlx_lm.server --model … --port 8080 — a plain OpenAI-compatible API) and oMLX (a native menu-bar app on :8000 that is a drop-in replacement for the OpenAI *and* Anthropic APIs, serving LLM/VLM/OCR/embedding/reranker models with continuous batching, a two-tier KV cache, and a web model-manager at /admin; install the .dmg from omlx.ai or brew install omlx). Point LibreChat, OpenCode, Claude Code or the Model Costs pane at either; both appear in the AI-harness model picker when running.
Run MLX with FrontierStack
FrontierStack lists MLX in its AI / LLMs catalog. Install or connect it from one place, then monitor its status, ports and certificate, secure it with the firewall and Malware Audit, and back it up.
Run it all from one Mac app.
FrontierStack installs, monitors and secures the whole stack — locally and across your fleet — from a single native macOS app.
Download FrontierStack