Rishab K Pattnaik

AI Research Engineer at Flam

Who I Am

I wasn’t supposed to end up here. Third year at BITS Pilani I picked up deep learning to fill a slot, got thoroughly nerd-sniped, and never really climbed back out.

Everything since comes down to one stubborn habit: taking something that wants a data centre and squeezing it into my mac and phone. That’s how three papers on reading X-rays happened: bones, fractures, chest films, all of it on the handset, no cloud involved.

These days I’m on the Avatar team at Flam, making faces talk in real time. GANs, diffusion, normalizing flows, and then the far less glamorous half where you fight WebRTC at 2am to shave off forty milliseconds.

If you’re building something in AI, world models especially, I’d like to hear about it. Hiring, starting something, or just want to argue about whether attention really was all we needed. The good conversations usually start there.

the practical bits

Now
AI Research Engineer at Flam, Bengaluru. Full time since July 2026, on the avatar stack. Started there as an AI R&D intern and stayed.
Studied
BE (Hons) in Electronics and Communication at BITS Pilani, graduated 26 July 2026. 3 publications came out of it.
Open to
Applied AI, ML and research engineer roles. Not hunting, but I’ll always take the conversation, and I’d move cities for the right one.
Reaching me
Email beats everything else: rishab27279@gmail.com. I actually read it, and I don’t do the let’s circle back thing. There’s a résumé if you need the formal version.

Experience

Flam

Jan 2026 – Present

Hamad Medical

May – Aug 2025

BITS Pilani

Aug 2024 – Dec 2025

IGCAR

May – Aug 2024

● active

AI Research Engineer

 Bengaluru, India  ·  Jan 2026 – Present  ·  AI R&D Intern until Jun 2026

  • Migrated the generation model from PyTorch to JAX. Static graphs and explicit RNG cut inference latency 30% on L40S, and made runs reproducible.
  • Built a half-body avatar POC by orchestrating LivePortrait with PEFT-adapted weights. Its output became synthetic training data for Flam's in-house avatar model.
  • Fine-tuned the renderer (not the generator) for an Indic use case. Lip blur on fast phonemes, texture sticking across frames, expressivity collapse — three failures, three different fixes.
  • Shipped real-time delivery end to end: WebRTC streaming, self-hosted GPU inference, streaming Indic TTS across 6 languages. Time-to-first-voice went from 3–4s to about 1s, with CUDA streams serving multiple concurrent clients off a single GPU.
  • Built detection-assisted data-curation tooling for the fine-tuning pipeline.

✓ completed

AI Research Intern

 Doha, Qatar (Remote)

  • Emergency research in the Dept. of Surgery, one of the Middle East's leading medical centres.
  • Multi-modal deep learning framework (PyTorch, Hugging Face) for automating emergency-department triage.
  • My contribution: the clinical-dataset analysis underlying the model.
  • The team's model hit 98% recall on Levels 1–3 and 93% top-2 accuracy on a custom dataset. Team's numbers, not mine.
  • Recall leads, not accuracy. Missing a Level 1 patient is not symmetric with over-triaging one.

✓ completed

Research Assistant — ECE Dept.

 BITS Pilani, Hyderabad

  • Three peer-reviewed outputs in 17 months, all asking the same question: choose the representation so the explanation comes free.
  • A lightweight Fourier Block Transformer for osteopenia and osteoporosis detection from X-ray, small enough to run on Android. Spectral blocks in place of spatial convolutions, so the features map to bone-texture bands a radiologist can name. IEEE Sensors Letters, first author.
  • A multi-frequency-aware CNN for bone fracture detection in muscle X-ray. Wavelet decomposition splits the signal into bands you can point at. Healthcare Technology Letters (Wiley), first author.
  • Deep representation learning for pneumonia and tuberculosis detection from chest X-rays. Elsevier book chapter, 2025.

✓ completed

CV Research Intern

 IGCAR, Kalpakkam

  • Adapted Meta's Segment Anything Model to camouflaged object detection on COD10K, a benchmark built specifically to break the boundary cues SAM depends on.
  • Measured the zero-shot baseline first, then built bounding-box-prompt fine-tuning against the gap it exposed, with a custom dataset class, dataloader and training engine.
  • The useful result was the negative one: vision foundation models don't transfer to specialised perception without adaptation.

Research & Publications

A Lightweight Fourier Block Transformer for Android-Based Edge‑Enabled Detection of Osteopenia and Osteoporosis

IEEE Sensors Letters · March 2026

Lightweight Fourier Block Transformer achieving 88.41% accuracy for real-time osteoporosis detection directly on Android devices using knee X-ray sensor images.

88.41% accuracy 60 imgs/min on Android TFLite edge deployment

Architecture: patch embedding + DFT block + dropout — no cloud required, fully on-device inference.

Rishab K Pattnaik  ·  Dr. R K Tripathy (BITS)  ·  Dr. R B Pachori (IIT Indore)

IEEE · 2026

Multi-Frequency Aware Deep Representation Learning for Bone Fracture Detection

Healthcare Technology Letters, 2025

Novel multiband-frequency aware network achieving 92.22% accuracy on bone fracture detection benchmarks.

92.22% accuracy AUC-ROC 0.959 6,229 MXR images

2D Wavelet Transform (LL/LH/HL/HH) + frozen EfficientNetV2B2 per subband — surpasses ResNet50, ViT, Swin Transformer.

Rishab K Pattnaik  ·  Dr. R K Tripathy (BITS)  ·  Dr. H Liu (Coventry)

IET / Wiley · 2025

Deep Representation Learning for Pneumonia & Tuberculosis Detection

Elsevier Book Chapter, 2025

Published in "Non-stationary and nonlinear data processing for automated computer-aided medical diagnosis".

87.08% accuracy Viral PN · Bacterial PN · TB IoT cloud deployment

Transfer learning blocks on chest X-ray images — classifies three respiratory disease categories with CAD-assisted diagnosis.

BITS Pilani · Published in Elsevier Non-stationary Data Processing volume

Elsevier · 2025

Featured Projects

U-Tube AI

Visual YouTube Agent

Multimodal AI that analyzes both audio and visual streams from YouTube videos.

EdgeSeg-AI

Research Paper Re-implementation

Memory-efficient image segmentation via sequential model loading.

Moody.AI

Multimodal Sentiment AI

Emotion recognition using DINOv2, Wav2Vec2, and DistilBERT for multimodal fusion.

OsteoDiagnosis.AI

Bone Health Diagnostics App

Android app for osteoporosis classification achieving 90% accuracy.

Expression.AI

Android App

Real-time facial expression recognition using TensorFlow Lite on Android.

Hand Gesture Identifier

Computer Vision

Real-time gesture recognition using FastViT achieving 97.5% accuracy.

AI Document Assistant

NLP / RAG

Document processing and Q&A pipeline using DeepSeek and Llama models.

Smart AI Checkers Bot

AI Game

Checkers game with a Minimax AI opponent with alpha-beta pruning.

Human Face Gender Detection

Neural Networks

Gender detection using InceptionV3 achieving 94.35% accuracy.

Technical Arsenal

Languages

Python
C
SQL
Git
Verilog

AI / ML Frameworks

PyTorch
TensorFlow
OpenCV
Scikit-learn
Keras
NumPy
Pandas
Matplotlib
Docker
LangChain
FastAPI
Android Studio
Streamlit

Specializations

Machine Learning
Deep Learning
Generative AI
LLMs

Education

Birla Institute of Technology and Science, Pilani

Bachelor of Engineering (Hons) in Electronics and Communication Engineering

Hyderabad, India  ·  2022 – 2026  ·  Graduated 26 July 2026

Relevant Coursework

  • Machine Learning for Electronics Engineer
  • Neural Network and Fuzzy Logic
  • Deep Learning
  • Digital Signal Processing
  • Signals and Systems
  • Probability and Statistics
  • Microprocessor and Interfacing
  • Linear Algebra

Blogs

MedMamba Explained

Deep Dive · Deep Learning · State-Space Models

The first Vision Mamba for Generalized Medical Image Classification — what it is, how it works, and why it matters.

Medium · 2024