π Hi, I'm Pavan Badempet
π Data & MLOps Platform Engineer | Big Data Architect
I am a Data Platform Engineer and MLOps Architect specializing in production-grade distributed lakehouses, high-throughput streaming pipelines, and scalable machine learning infrastructure.
ποΈ Core Architectures: Medallion Lakehouses (Delta Lake, Databricks), Real-Time Streaming (Kafka, PySpark), and Workflow Orchestration (Apache Airflow).π€ AI & Inference: Tabular Foundation Models, Recommendation Ensembles (SASRec, LightGCN), Vector Search (pgvector), and Local LLM RAG pipelines.π₯ Industry Domains: Clinical Health Informatics (FHIR R4, OMOP CDM v5.4, ABDM compliance), Capital Markets, and Media Recommendation Platforms.
π Technical Superpowers & Stack
π Featured Flagship Production Platforms
π₯ AI-Healthcare-System
Enterprise AI Healthcare Lakehouse & Clinical Intelligence Platform
- Architecture: PySpark Medallion Lakehouse, Apache Airflow pipelines, FHIR R4 / OMOP CDM v5.4 compliance, and HIPAA-ready FastAPI backend.
- ML & RAG: TabICLv2 Tabular Foundation Models, calibrated CatBoost/XGBoost ensembles with 95% Conformal Prediction sets, 10-year multi-organ digital twin simulator, and local Ollama LangGraph multi-agent RAG.
- Live Links:
π€ Hugging Face Live Interactive Space Β·π€ Model Hub (16 Weights) Β·π¦ GitLab Repository
π¬ AI-Recommendation-System
Real-Time AI Media Recommendation Engine & Unified Data Intelligence Platform
- Scale: Ingests and processes 21M+ real records (1M+ TMDB movies and 20M+ MovieLens ratings) across a Databricks Serverless Medallion Lakehouse.
- Deep Learning & Search: SASRec sequential transformers, LightGCN graph embeddings, and a 10-shard Neon Serverless
pgvectorHNSW cluster (<5ms query latency). - Live Links:
π Live Portfolio Showcase Β·π€ Hugging Face Space UI Β·π¦ GitLab Repository
π GitHub Activity & Metrics
π Connect & Professional Links
π Career Keywords & Technical Index (Search Engine Optimization)
Core Specializations: Lead Data Engineer, Senior Data Architect, Big Data Engineer Portfolio, MLOps Engineer, Lakehouse Architect, Python Developer, PySpark Specialist, Distributed Systems Engineer. Distributed Computing & Lakehouse: Apache Spark, PySpark Streaming, Delta Lake, Apache Iceberg, Apache Airflow, Databricks Medallion Lakehouse, Data Vault 2.0, SCD Type 2, Liquid Clustering, Z-Order Optimization, Great Expectations. Machine Learning & Search: Recommendation Engines, SASRec, LightGCN, Graph Neural Networks, PyTorch, pgvector HNSW Indexing, Vector Databases, Conformal Prediction, Hugging Face Spaces, Ollama Local Inference, LangGraph Multi-Agent RAG. Cloud & Infrastructure: Amazon Web Services (AWS EMR, S3, Glue, Athena, RDS, ECS), Docker Containerization, Kubernetes Orchestration, CI/CD GitHub Actions, PostgreSQL, Redis Streaming.
