I build production LLM systems — agentic pipelines, RAG applications, and AI backend infrastructure — as an AI Engineer at InterCraft. My focus is reliability: full observability stacks, structured evaluation, and adversarial red-teaming (OWASP LLM Top 10, EU AI Act) before anything ships.
- 🧠 Currently building agentic systems on LangGraph + FastAPI, instrumented end-to-end with OpenTelemetry & Grafana
- 🛡️ Red-teamed 700+ adversarial prompts across two production LLM systems
- 🎓 BSCS, CGPA 3.87 — Arid Agriculture University, Rawalpindi
AI Engineer — InterCraft · Sep 2025 – Present
- Architected the Alfaisal University Helpdesk AI System — full email-to-ticket lifecycle on LangGraph, FastAPI, and Celery/Redis, integrated with a Laravel backend via Microsoft Graph webhooks
- Led a red-team security assessment (OWASP LLM Top 10, EU AI Act, ISO/IEC 42001) — 207 adversarial tests, 6 vulnerabilities identified
- Instrumented the stack with OpenTelemetry and built Grafana dashboards (Loki, Tempo, Prometheus) for production observability
- Shipped a University Policy RAG Chatbot (LangChain, Qdrant) and an AI auto-grading pipeline with structured CSV export
Machine Learning Intern — Tensor Labs · Mar 2025 – Jun 2025
- Optimized production AI pipelines (LangChain, HuggingFace, Google ADK), improving efficiency by 30%
- Built LLM-based agents for task automation, retrieval, and code execution
Evaluation, Safety & Observability
Production agentic RAG system for semantic search across Quran, Hadith, and Fatwa corpora. RAG Fusion with Reciprocal Rank Fusion, Redis semantic cache, Firebase OAuth, and automatic LLM fallback (AWS Bedrock Mistral → Groq LLaMA 3.3). Red-teamed with Promptfoo — 539 adversarial tests, 89% defense rate against jailbreaks.
LangGraph FastAPI Qdrant Redis AWS Bedrock Promptfoo
Real-time voice agent handling STT → LLM → TTS end-to-end, containerized and deployed live on my portfolio.
LiveKit Deepgram ElevenLabs Groq AWS Docker
Also built: an autonomous coding agent with multi-agent orchestration (Google ADK + LangGraph), a DenseNet201 skin-disease classifier (99.15% accuracy), and a heart-disease risk model (89% accuracy).
- LangChain Essentials — LangChain Academy
- Supervised ML: Regression & Classification — DeepLearning.AI / Stanford
- AWS AI Practitioner — LearnKartS (Coursera)
- Responsible AI with AWS Security & Governance — LearnKartS
