Analytics Engineer · Data Analyst · Healthcare

I build data models
teams can trust.

I turn claims, EMR, and CMS data into decisions for actuarial, clinical, and executive teams — with pipelines that are tested, documented, and built to hold up. Currently at Humana.

source Raw systems 25+ feeds
staging Clean & cast type-safe
marts Star schemas facts · dims
decisions Trusted output dashboards
0Medicare Advantage members in the data I model
0faster daily pipeline runtime
0HCC risk-score accuracy gain (Actuarial-validated)
0tested dbt projects on public data (open source)

-- selected work

Things I've built end to end.

Analytics engineering, full-stack products, and the pipelines that ship them. Each links to source or a live site.

Analytics engineering

13 tests · passing code ↗

CMS Hospital Readmissions Analytics

The public CMS HRRP dataset modeled into a tested star schema — null-safe casting for suppressed values, surrogate keys at facility × measure grain, and a state-level excess-readmission-ratio mart.

  • Star schema with surrogate keys at facility × measure grain; null-safe casting for suppressed values.
  • 13 dbt tests — unique, not_null, relationships, accepted_values.
  • State-level excess-readmission-ratio mart, built for BI reporting.
  • dbt
  • SQL
  • DuckDB
  • Star schema
20 tests · passing code ↗

Synthea Health Analytics

An end-to-end pipeline — staging → intermediate → marts — over synthetic patient data. Star schema and a chronic-disease prevalence mart segmented by age band and gender. No PHI; reproducible from raw CSVs.

  • staging → intermediate → marts over synthetic patient, encounter, and condition data.
  • Chronic-disease prevalence mart segmented by age band and gender.
  • 20 dbt tests + dbt Docs; reproducible from raw CSVs, no PHI.
  • dbt
  • SQL
  • DuckDB
  • Python

Data & cloud engineering

6 AWS services · streaming

Elastic GPU Gating — Serverless AWS Pipeline

Independent research project — an event-driven AWS pipeline that gates GPU activation on a coverage-constrained safety score instead of workload volume.

  • Event-driven flow: Kinesis → Lambda → DynamoDB → SQS → EC2 → SNS.
  • Application-level exactly-once via idempotent DynamoDB writes.
  • 33,475 records replayed — 89% per-node energy reduction (emulation), P99 2.9 s.
  • Reproducible: notebooks, Lambda/worker code, one-command deploy.
  • AWS Lambda
  • Kinesis
  • DynamoDB
  • SQS / SNS
  • Python
  • Streaming

Production analytics platforms

Live visit ↗

Mainaka

Analytics Platform Engineer · India. A production analytics platform I built and operate over the full applicant journey — event instrumentation, operational metrics, and executive reporting on live user activity.

  • GA4, Google Tag Manager, and Microsoft Clarity across a React SPA, with custom dataLayer events for funnel tracking (registration, mock interviews, sign-ups, engagement).
  • Fixed a production pipeline bug — redirected outcome submissions from a deprecated table to the canonical dataset, restoring KPI accuracy.
  • Node/Express APIs over SQLite/PostgreSQL feeding operational dashboards; automated zero-downtime deploys with monitoring, backup, and rollback.
  • GA4
  • GTM
  • React
  • Node.js
  • PostgreSQL
  • Funnel analysis
Live visit ↗

GlobalGoGateway

Analytics Platform Developer · India. A production workflow-and-analytics platform I built and operate for customer-lifecycle management, operational processes, and business-performance reporting.

  • PostgreSQL operational database (Customer360, Case360, workflow, tasks, documents, payments) with event-driven timeline tables modeled for analytics.
  • Operational dashboards tracking pipeline stages, case aging, approval outcomes, consultant productivity, and business KPIs.
  • ~100 real customer cases curated into a historical operational dataset — case aging, approval outcomes, and lifecycle events — for workflow and business analytics.
  • PostgreSQL
  • SQL
  • Event modeling
  • KPI dashboards
  • Chart.js

-- how i build

Layered, tested, and fast.

The same shape every time: raw sources refined through tested layers into star schemas teams can query with confidence.

Dimensional model · star schema
Pipeline runtime · re-engineered
Before 3.0 h
After 1.8 h

Incremental models + partition pruning — ~40% faster daily runtime.

Data quality · every layer
  • ✓ unique
  • ✓ not_null
  • ✓ accepted_values
  • ✓ relationships
  • ✓ freshness
  • ✓ range

Generic + custom tests catch data anomalies early — before they reach a dashboard.

-- experience

Where I've done the work.

  1. Mar 2024 — Present Arlington, VA

    Analytics Engineer

    Humana · Service Fund (value-based care)

    • Build and maintain the dbt + Snowflake transformation layer (staging → intermediate → marts) over claims, member eligibility, and clinical data spanning 5M+ Medicare Advantage members.
    • Cut daily pipeline runtime from 3h to 1.8h (~40%) by introducing incremental models and partition pruning.
    • Partnered with Actuarial to correct misclassified CMS-HCC codes, improving risk-score accuracy by 8% (Actuarial-validated).
    • Migrated 30+ legacy SSIS packages to dbt + Airflow; added dbt tests that catch data anomalies early, reducing downstream issues. Built utilization and provider-performance dashboards in Tableau.
  2. Sep 2020 — Nov 2022 Hyderabad, India

    Data Analyst — Analytics Engineering track

    Apollo Hospitals · Connected Care Command Centre

    • Built Python ETL pipelines integrating EMR, operational, and demographic data from 25+ source tables into Azure SQL and Snowflake; scheduled and monitored with Airflow.
    • Built a length-of-stay tracker and a daily bed-occupancy dashboard in Power BI, used by hospital administrators; contributed to a ~6% reduction in average length of stay over 12 months.
    • Automated weekly performance reports in Python, replacing manual Excel work (~10 hours/week saved) and cutting data-reconciliation time from 4h to 1.5h.
    • Mentored 2 junior analysts to independence on SQL and dashboard development.

-- stack

What I work with.

Modeling & SQL

  • dbt (Cloud & Core)
  • SQL
  • Dimensional modeling / star schema
  • Snowflake
  • Azure SQL
  • PostgreSQL
  • DuckDB

Pipelines & cloud

  • Python (Pandas)
  • Apache Airflow
  • AWS (EC2, S3, IAM)
  • Git / GitHub
  • CI/CD

BI & domain

  • Tableau
  • Power BI
  • CMS-HCC risk adjustment
  • HRRP
  • value-based care
  • claims & EMR data

Also

  • Node.js
  • Express
  • React
  • LLM integration

Education

M.S., Business Analytics

Grand Canyon University, Phoenix, AZ

Jan 2023 — Jul 2024

B.Tech, Electronics & Communication Engineering

V. R. Siddhartha Engineering College, India

Jun 2017 — Mar 2021

Certifications

  • dbt Analytics Engineering — dbt Labs, 2024
  • AWS Certified DevOps Engineer – Professional, 2024
  • AWS Certified Solutions Architect – Associate, 2024

-- get in touch

Have a data problem worth
building for? Let's talk.