About Projects Experience FAQ Contact

Hi, I'm

NAYEEM
FARDIN

AI-Native Systems Engineer

Scroll
AWS• KUBERNETES• FASTAPI• REACT• KAFKA• RABBITMQ• TERRAFORM• POSTGRESQL• GO• AWS• KUBERNETES• FASTAPI• REACT• KAFKA• RABBITMQ• TERRAFORM• POSTGRESQL• GO•

ABOUT.

I'm a Computer Science undergrad at Islamic University of Technology who builds AI-native systems — computer vision, RAG, and LLM pipelines — on production-grade backend foundations designed to scale and break gracefully. My focus is shipping AI features with real engineering rigor: async queues, observability, reliability engineering, and cost-conscious, cloud-native architecture.

Education

B.Sc. in Computer Science & Engineering

Islamic University of Technology (IUT)

Expected Sep 2026

IUT Logo

Tech Stack

Languages
Python Go Dart Kotlin JavaScript C++ Java
Web & API
React Node Express FastAPI
Mobile
Flutter Android
AI & ML
LangChain PyTorch RAG OpenAI
Cloud & Ops
AWS Docker Kubernetes Terraform Linux Git
Sys & Tools
Kafka RabbitMQ Celery Grafana ArgoCD Helm
Databases
PostgreSQL Redis InfluxDB Pinecone MongoDB

PROJECTS.

RetailOS Lite

RetailOS Lite supervisor overview dashboard preview
Next.js 15 TypeScript FastAPI BullMQ YOLO Prisma Modal GPU Pinecone Prometheus Grafana
  • AI-native retail execution platform featuring async YOLO shelf analysis, multi-signal fraud detection, and governed outlet master data.
  • Orchestrates an async BullMQ pipeline executing YOLO + LLM reasoning, complete with Prometheus queue metric collection and DLQ replay CLI.
  • Features a 5-signal calibrated fraud engine (SHA-256, dHash, GPS mismatch, EXIF analysis) that runs early to save expensive GPU resources.
  • Implements weighted outlet similarity matching (pg_trgm + geo prefilter) with three-tier resolution and non-destructive duplicate merging.
  • Designed a two-tier RAG operational assistant combining exact DB queries (via Prisma) with Pinecone semantic search over visit reports.

DeliveryLens

Next.js 16 TypeScript Neon Postgres pgvector TypeORM Zod Cohere Rerank Recharts
  • AI delivery-intelligence assistant built during an AI engineering internship — answers project status, blockers and team tracking from a local mirror of GitHub and a task tracker.
  • The model writes the text but never carries a number: 16 typed tools (Zod) run SQL, and a view builder renders charts from the same result JSON with no model in the path.
  • Hybrid retrieval fusing pgvector/HNSW and Postgres full-text by reciprocal rank, then a Cohere cross-encoder rerank, with the per-document cap applied after scoring.
  • Reranker failure degrades loudly to a diversity fallback that needs no extra API call, and records which path ran so a weaker answer is flagged rather than shipped silently.
  • Aggregates carry their own denominators and caveats; incomparable roles are ranked in separate groups, and zero-activity accounts are listed apart from ranked ones.
  • Branch-aware incremental GitHub sync with per-repo watermarks, built after a 65-branch repo exhausted the hourly rate limit.
Deskemy course library home screen preview
Rust Slint libmpv OpenGL Skia SQLite FTS5 minisign
  • Shipped Windows desktop course player with 52 GitHub stars, 343 downloads across 7 releases, and a community-maintained Chinese translation.
  • 2.0 is a full interface rebuild: from Tauri and WebView2 to a native Slint UI, removing the Windows-only compositor that was blocking Linux and macOS.
  • libmpv's render API draws straight into the UI's own OpenGL context, so the controls, subtitles and panels over the video are just ordinary elements on every platform.
  • Rewrote without a flag day: a UI-agnostic core crate shared by the old and new apps while the old one kept shipping.
  • Upgrades 1.x installs in place without losing a library; updates are opt-in and signed, and the app rejects any that fail verification.
  • Also: an always-on-top mini player, a two-phase importer, rename-safe progress, and SQLite FTS5 search over spoken subtitle text.

HomeLab

Homelab dashboard preview showing host stats, DNS and proxy metrics, and offsite backup health
Podman Quadlets systemd SELinux CoreDNS Caddy Tailscale rclone
  • Two hosts running 19 rootless podman containers as declarative systemd Quadlets — no root daemon, no compose, everything up at boot with nobody logged in.
  • Split-horizon DNS serves one set of .home names to two audiences: a LAN resolver and a tailnet-only resolver both pointing at the same reverse proxy.
  • Images pinned by @sha256: digest so a pull can never silently change what runs; SELinux stays enforcing with per-path relabel decisions.
  • Offsite backups verified by restoring them — an encrypted remote exposes no usable hashes, so a weekly job pulls back 15 random files, decrypts them, and checksums against the original.
  • Storage placement dictated by a shingled (SMR) disk: write-heavy and latency-sensitive data on SSD, only large sequential data on the HDD.
  • Incident log with real postmortems — including one where three broken discovery paths had three different causes, and the fix then broke three more things.
Python YOLOv11 AWS EKS FastAPI Docker ONNX Prometheus ArgoCD Helm
  • Fine-tuned and deployed YOLOv11L on AWS EKS as a microserviced inference service for dense, multi-class retail product detection.
  • Compiled the model into an in-process ONNX Runtime engine with INT8 quantization + right-sized CPU bin-packing (200m/pod) for sub-500ms detection entirely on CPU — no GPU required.
  • 4× CPU inference speedup and 70–90% cost savings via INT8, AMX, and Spot node diversification across m7i-flex, c7i-flex, and t3.small.
  • GitOps progressive delivery: ArgoCD Image Updater polls ECR, Argo Rollouts runs weighted canaries through NGINX Ingress with automated smoke tests before promotion.
  • Spot resilience enforced with Pod Disruption Budgets and native 2-minute interruption handling for zero-downtime operation on volatile hardware.
  • DevSecOps gates in GitHub Actions (OIDC/STS) — ruff, pip-audit, and Trivy CVE scans guard the ECR boundary; Loki/Prometheus/Grafana for observability.
Go Python Celery RabbitMQ Terraform AWS K3s Redis
  • Migrated the HTTP edge from FastAPI to Go, preserving full contract parity with the existing Celery worker runtime and Alembic-managed schema.
  • Architected a dual-node K3s cluster on AWS (On-Demand control + Spot data plane) with full Terraform IaC and automated SSM reconciliation for async image processing queued through RabbitMQ.
  • Optimized cost with EC2 Spot and graceful eviction via AWS Node Termination Handler.
  • Resilience under load with Redis idempotency keys, pybreaker circuit breakers, and k6 stress tests.
  • Zero-Trust GitHub Actions (OIDC) CI/CD; telemetry and traces to Grafana LGTM via Alloy.
  • Decoupled ML inference from fast conversions via dedicated queues to prevent worker starvation.
Flutter Kotlin CameraX ML Kit AlarmManager GeofencingClient MapLibre
  • Flutter/Native split: Flutter UI shell backed by a Kotlin alarm engine that owns scheduling, ringing, recovery, and dismissal authority.
  • Location alarms via GeofencingClient with hybrid geofence + passive approach-assist and 10-state health model per alarm.
  • Mission-based dismissal with native inactivity enforcement — math, steps (TYPE_STEP_DETECTOR), and QR (CameraX + ML Kit) missions.
  • Direct-boot persistence and reboot recovery before first unlock via device-protected storage and LOCKED_BOOT_COMPLETED.
  • 28-finding security/performance/reliability audit driving a dedicated hardening sprint. Macrobenchmark + Perfetto performance tooling.

More Projects.

StatusMonitor

FastAPI Kafka InfluxDB React 19

Distributed monitoring platform — FastAPI microservices, a Kafka ingestion pipeline, tiered InfluxDB retention, and real-time Telegram alerts.

Crawler RAG DocHelper

LangChain Pinecone Gemini 2.5 Cohere

Hybrid-search RAG over docs — semantic + BM25 retrieval in Pinecone, Cohere rerank, and Gemini 2.5 generation with query-time alpha tuning.

interviewPrepper

LangGraph FastAPI Multi-LLM Firecrawl

AI 30-day learning-plan generator — LangGraph agent orchestration, cross-LLM routing (OpenAI/Gemini/OpenRouter), and Firecrawl + Tavily retrieval.

RoutineMaker

FastAPI Docker Nginx PostgreSQL

Microservices routine manager — split Auth/Core FastAPI services behind an Nginx gateway, Argon2 + JWT security, and automated PDF export.

EXPERIENCE.

Intelligent Machines Logo

AI Native Engineering Intern

Intelligent Machines

June 2026 — August 2026
  • Built production-grade AI-native systems at a company shipping 59 production systems across 16 enterprise clients in 6 countries — including Unilever, bKash, and Banglalink.
  • Worked across the AI/ML engineering stack spanning computer vision, document intelligence, and enterprise integration pipelines for trade marketing and distribution platforms.
Backdoor Private Limited Logo

CyberSecurity Internship

Backdoor Private Limited

October 2025
  • Digital forensics analysis to recover sensitive deleted artifacts using FTK Imager and Autopsy.
  • Security assessments on Android applications via MobSF and web platforms using Burp Suite and SQLMap.

FAQ.

What does "AI-Native Systems Engineer" actually mean?

It means I care about the systems underneath the AI, not just the prompt on top of it. Wrapping a model in an app is the easy part. The engineering is in the queues that keep inference off the request path, the retrieval that has to return the right passage, the boundaries that stop a model touching things it should not, the metrics that tell you it is degrading, and the cost of every GPU second. In DeliveryLens the model writes the prose but never carries a number, because a small model once mis-transcribed a year out of a correct result. That rule is what the title means to me.

Why do you build systems instead of just AI demos?

Because a demo gets to assume the model is fast, available, and right. Real systems do not. The interesting questions all start where the demo ends: what does the user see while inference is running, what happens when the provider rate-limits you, how do you tell a wrong answer from a slow one, and what does this cost at a hundred times the traffic. I gravitated to backend and infrastructure because that is where those answers live. AI made the questions more interesting, not different.

How do you approach building AI systems that can fail safely?

Four habits, mostly learned by getting them wrong first. Keep the model off the critical path — in RetailOS Lite a visit submission enqueues work and returns, so a slow model never blocks a field rep. Give the model narrow boundaries — YOLO produces grounded detections, an LLM reasons about them, and typed tools running SQL produce every number, so nothing measured depends on the model remembering it. Make degradation loud — when the Cohere reranker times out, retrieval falls back to a diversity pick and records which path ran, so a weaker answer is flagged rather than shipped silently. Spend cheap checks before expensive ones — the fraud engine runs five cheap signals before anything reaches a GPU.

What happens when one of your systems breaks?

Ideally I can answer "what broke, when, and for whom" without guessing. That means queues with dead letters and a replay CLI that defaults to a dry run, correlation IDs threaded across services, Prometheus and Grafana with Loki for logs, and results that carry provenance — when each source last synced, what was excluded, which tasks could not be judged at all. My homelab is where I practise the unglamorous half: 19 containers with an incident log and real postmortems, and a weekly job that restores a random sample of the backup and checksums it, because a backup you have never restored is a guess.

Why are your projects so infrastructure-heavy?

Because that is where things actually break, and I find that part genuinely interesting rather than a chore. Spot instances need graceful eviction or you lose work. A 65-branch repository will exhaust GitHub's hourly rate limit unless the sync is incremental and watermarked. Container images pinned by tag can silently change what runs, so mine are pinned by @sha256: digest. A shingled drive stalls for seconds under random writes, so write-heavy data belongs on the SSD and only sequential bulk goes on the HDD. None of that shows up in a feature list, and all of it decides whether the thing stays up.

What have you built that you're actually proud of?

RetailOS Lite — an AI-native retail execution platform with async shelf analysis, a five-signal fraud engine that runs before the GPU, and a RAG assistant that queries exact data before it reaches for vectors. DeliveryLens — a delivery-intelligence assistant where 16 typed tools run SQL so no measured number is ever written by a model. Deskemy — the one with real users. ShelfWatch — dense retail detection on AWS EKS, 4× faster CPU inference and 70–90% cheaper through INT8 quantization and Spot diversification. PixTools — a Go HTTP edge over Celery workers on a Terraform-provisioned K3s cluster, with circuit breakers and idempotency keys for when things go wrong.

What's the weirdest thing you've built?

Deskemy, an offline course player for Windows, because I have now solved the same problem in it three times. It plays through libmpv rather than an HTML5 video element, so it handles whatever codecs a downloaded course ships with — but then something has to draw the controls over that video. The first version put the video in a separate window above a web view, and the picture lagged the layout on every resize. Version 1.x composited the video beneath a transparent web view, which fixed it but only on Windows. 2.0 rebuilt the whole interface in Slint so the video renders into the UI's own OpenGL context and the problem simply stops existing. It has 52 stars, 343 downloads, and a community-maintained Chinese translation, which I did not expect.

What are you currently learning?

Operating things properly, rather than only building them. Most of that happens in a two-machine homelab running 19 rootless podman containers as systemd Quadlets, which has taught me more about SELinux labelling, split-horizon DNS, mount ordering, and digest-pinned images than any tutorial would have. I am working toward platform and reliability engineering — deeper Kubernetes than I have needed so far, and the distributed-systems fundamentals behind consensus, backpressure, and failure detection. I would rather say that plainly than pretend I have finished.

What kind of problems do you want to work on?

AI infrastructure and platform engineering — the layer that makes model-backed features dependable rather than impressive once. Inference serving and the queues around it, retrieval systems that have to be right and not just plausible, and the observability that tells you which of the two you have. More broadly: backend and distributed systems where correctness under failure is the actual requirement, and developer platforms where one person's work removes friction for everyone else. Cost-aware infrastructure interests me too, because a system nobody can afford to run is not finished.

How can we work together?

I graduate in September 2026 and I am looking for engineering roles in AI infrastructure, backend, or platform work — new-grad positions and internships both. I am also happy to talk about project collaborations, or about any of the systems on this site if something here is useful to you. The fastest route is email at nayeemfardin00@gmail.com; my code is on github.com/NFRohan and I am on LinkedIn.

LET'S BUILD
SOMETHING.

Got a project idea, internship opportunity, or just want to chat? Hit me up.