LocalForge Router
A self-hosted LLM control plane with GGUF speculative routing, dynamic ML dispatch, and VRAM-aware model lifecycle management.
Muhammad Ali Nasir is an ML Systems Engineer with a focus on high-throughput architectures, causal inference, and compact intelligence. I help engineering teams architect sub-5ms routing pipelines, benchmark verified research models, and ship autonomous systems that perform in production.
Explore seven production-grade machine learning systems, from sub-5ms speculative routers to causal pricing engines and parameter-efficient transformers.
A self-hosted LLM control plane with GGUF speculative routing, dynamic ML dispatch, and VRAM-aware model lifecycle management.
Multi-tenant pricing intelligence platform forecasting demand, estimating causal price elasticity, and calculating bounded profit frontiers.
Four-agent deliberation over a Neo4j scientific paper graph. Dynamic tool loading cuts context overhead while maintaining evidence fidelity.
Parameter-efficient transformer architecture fine-tuned on 24k Urdu comments with dynamic layer pruning and knowledge distillation.
Published command-line intelligence tool for chatting with multi-repo codebases using tree-sitter AST syntax parsing and CrewAI agents.
GIKI research project combining Extra Trees and LightGBM solar irradiance forecasting with thermodynamic cooling optimization across Pakistan.
State-space DP optimizer balancing vehicle fuel capacity, safety reserves, and detour costs with real-time station pricing for US routes.
A lightweight transformer for four-class Urdu hope-speech classification, built around a 24,124-sample dataset and interpretable feature analysis.
Abstract: Online hope speech classification is critical for promoting positive engagement in low-resource languages. We present LightUHope, a parameter-efficient transformer architecture fine-tuned on 24,124 annotated Urdu social media comments. By deploying dynamic layer pruning and knowledge distillation, LightUHope achieves a 0.92 macro F1 score using only 3.2M parameters—a 97% reduction compared to mBERT while maintaining full zero-shot generalization capabilities across regional dialects.
Pakistan Institute of Engineering and Applied Sciences (PIEAS), Islamabad
Whether your system needs sub-5ms speculative routing, high-throughput causal intelligence, or custom LLM fine-tuning pipelines — I'm ready to architect and ship it.