Data Analyst / Data Engineer / Data Scientist

Binju Karki

I build pipelines, models, and dashboards that sit underneath decisions that matter to real people: I don't ship a number I can't trace back to its source.

Pursuing an MS in Data Science & Analytics at Grand Valley State University (Dec 2026). Currently building the ETL pipeline and Power BI platform behind reporting for 78 charter schools. Previously a data engineer in healthcare analytics, and now building agentic AI tools that ground every answer in real data instead of guesswork.

SQLPythonR Power BIETLDatabricks
78Charter schools, current project
3.9 / 4.0Graduate GPA
2Deployed AI agents, live in production
2+ yrsProfessional analytics & engineering
Binju Karki
01

Ground truth

Every role I've had has had the same catch: get the data wrong, and it's not an abstraction that breaks: it's a real person on the other end of it.

At Cedar Gate, that meant HIPAA-bound healthcare records. At GVSU's Charter Schools Office, it's FERPA-bound student records for 78 schools: data that gets validated and suppressed before it ever reaches a dashboard, because the person behind each row is a kid. Before that, as a Study Abroad Peer Advisor, the "data" was a student in front of me deciding whether to spend a semester on the other side of the world.

That's the thread running under everything I build now, including the AI agents: I don't let a model, a report, or a dashboard say something it can't back up.

2023
Healthcare Data, Cedar Gate · HIPAA & PHI
2025
Study Abroad advising · Real students, real stakes
2025–26
Student Data, GVSU · FERPA & PII, 78 schools
2026
AI Agents · Grounded, not guessing
02

Projects & results

What I've built, what it does, and what came out of it: from a BI platform in progress for 78 schools to deployed AI agents and statistical research.

R · Power BI · DAX · Power Query · Star Schema

K-12 Data & Reporting Platform

Building now

I own the data pipeline behind GVSU's Charter Schools Office reporting: I've already cleaned, integrated, and validated K-12 data across all 78 schools, and I'm now building the interactive Power BI dashboard that unifies compliance, financial, academic achievement, academic growth, enrollment, and peer-benchmarking reporting into one place per school.

Completed: data cleaning, integration, and validation across enrollment, assessment, and financial sources for all 78 schools. In progress: a unified interactive dashboard covering all six reporting domains for all 78 schools at once.
78
schools, cleaned data & dashboard in one place
6
reporting domains unified: compliance, financial, achievement, growth, enrollment, peer benchmarking
R + Power BI
cleaning pipeline & dashboard build

Platform architecture

Sources Enrollment (SIS) Assessment data Financial indicators Clean & Validate (done) R data cleaning integration · QA checks validation rules Model (building now) Star schema Power Query + DAX data suppression Serve (building now) Power BI dashboards FERPA / ADA compliant ● in progress Consumers 78 schools State compliance Public reporting 2025: cleaning & validation completed for all 78 schools · 2026: building the unified dashboard (model + serve) now

Diagram reflects the pipeline architecture I built and am building; individual student-level data is FERPA-protected and never shown.

Agentic AI · Streamlit · OpenAI

Research Synthesis Agent

Live

An agentic research assistant that reads uploaded papers or abstracts and produces a structured literature-review synthesis (themes, methodologies, conflicting findings, and research gaps) instead of a generic summary.

Result: a working, publicly deployed agent that autonomously decides when to pull outside research via tool-calling, rather than only summarizing what's given to it.
1
custom retrieval tool
6
structured output fields
GPT-4o-mini
model
Agentic AI · FastAPI · Streamlit · OpenAI

MentorMate · AI Study Companion

Live

An agentic study assistant for graduate coursework that decides, per question, whether it needs a term definition, course-note context, or a study plan before it answers.

Result: 100% correct autonomous tool selection across 8 evaluation cases: the agent reliably picks the right tool instead of guessing or hallucinating an answer.
3
orchestrated tools
100%
tool-selection accuracy, 8 eval cases
2
deployment paths (FastAPI + Streamlit)
R · tidymodels · Multinomial / LDA / QDA / Poisson

Family Involvement & Student Grades

A classification study on the NCES Parent & Family Involvement in Education survey, predicting a student's usual letter grade from parental engagement and household background.

Result: multinomial logistic regression was the best of 3 models tested, but with a majority-class baseline of 56.5% (most students get an "A"), its 58.9% CV accuracy is a modest lift, not a strong classifier. The more useful finding: family involvement (parent-teacher conferences, activity attendance) was a statistically significant predictor of grade category even after controlling for household income.
13,417
households analyzed
3
models compared, 10-fold CV
58.9%
best CV accuracy
R · Monte Carlo Inference · ggplot2 · Leaflet

U.S. State-Level Census & Election Analysis

Merged state-level census and presidential election data (2008–2016) to test whether income, education, and housing cost differ between states won by each party.

My contribution: ran a 5,000-iteration permutation test showing education levels differ significantly between Democrat- and Republican-won states, and a 10,000-sample bootstrap producing a 95% CI of $55.6K–$57.6K on national median income.
5,000
permutation iterations
10,000
bootstrap samples
3
election cycles compared
R · tidymodels · 8 Modeling Techniques

STA 631 Statistical Modeling Portfolio

A self-directed modeling portfolio built for STA 631 (Statistical Modeling), covering 8 techniques end-to-end in R using the tidymodels framework: multiple linear regression, polynomial regression, interaction models, multinomial logistic regression, LDA, QDA, Poisson regression, and ridge/lasso penalized regression, applied to the Auto and Diamonds datasets with full diagnostics and cross-validation.

Result: compared 3 classification approaches for predicting a car's origin (up to 78.8% test accuracy), and 2 penalized regression approaches for predicting mpg, with lasso outperforming ridge (RMSE 3.38 vs. 3.52). An interaction model on the Diamonds dataset reached R² = 0.998.
8
modeling techniques, one portfolio
0.998
R² on best-fit interaction model
78.8%
best classification test accuracy (QDA)
D3.js · Observable · Interactive Visualization

A Visual Analytics Approach to Wine Quality

Ten interactive D3.js visualizations exploring what physicochemical properties drive wine quality, built for wine producers, sommeliers, and enthusiasts to explore directly.

Result: identified alcohol content and volatile acidity as the two strongest quality predictors, and shipped a working blend-simulator so users can test the finding themselves.
6,497
wines in dataset
10
interactive visualizations
11
physicochemical features
03

Skills

Languages & Data Engineering

PythonRSQLDAX Power QueryPySparkDatabricks Amazon RedshiftMySQLAmazon S3ETL Pipelines

Visualization & Analysis

Power BIExcelCanva Exploratory Data AnalysisStatistical Data AnalysisStar Schema Modeling

Governance & Compliance

HIPAAFERPAPII Handling Data GovernanceData ValidationADA Compliance

Applied ML & AI

Scikit-learntidymodelsOpenAI API Function CallingGrounding / RAGStreamlit
Critical ThinkingCross-functional Collaboration Interpersonal SkillsAdaptability Teamwork
04

How I work

01

Validate before I visualize

Every dashboard I ship has been through a data quality and validation pass first. A clean-looking chart built on unvalidated data is worse than no chart at all.

02

Ground before I generate

Whether it's a compliance report or an AI agent's answer, I don't let output get ahead of its source. My agents cite what they used; my reports document where the numbers came from.

03

Build for the person who has to trust it

A school administrator, a healthcare client, a student deciding on a semester abroad: the end user isn't a data person. If they can't act on it, the analysis isn't finished yet.

05

Code samples

Real snippets pulled straight from my own scripts, not illustrative pseudocode.

Census & Election analysis (R) · permutation test

# Two-sample t-test on observed groups
ttest_result <- t.test(formula = bach_higher ~ party_clean,
      data = trial_data, alternative = "two.sided")

n_permutations <- 5000
permutation_statistics <- vector(length = n_permutations)

for(p in 1:n_permutations) {
  permutation_statistics[p] <- t.test(formula = bach_higher ~ party_clean,
      data = trial_data |> mutate(party_clean = sample(party_clean))) |>
      broom::tidy() |> pull(statistic)
}

Census & Election analysis (R) · bootstrap CI

income_data <- state_census |>
  filter(!is.na(median_income)) |>
  pull(median_income)

B <- 10000
boot_medians <- replicate(B, {
  sample(income_data, size = length(income_data),
         replace = TRUE) |> median(na.rm = TRUE)
})

boot_se <- sd(boot_medians)
boot_ci <- quantile(boot_medians,
      probs = c(0.025, 0.975))

Family Involvement & Grades (R / tidymodels) · model workflow

multi_spec <- multinom_reg(mode = "classification") %>%
  set_engine("nnet")

multi_wf <- workflow() %>%
  add_recipe(grade_rec) %>%
  add_model(multi_spec)

set.seed(123)
multi_res <- multi_wf %>%
  fit_resamples(
    resamples = cv_folds,
    metrics = metric_set(accuracy)
  )

STA 631 Modeling Portfolio (R / tidymodels) · LDA spec

lda_spec <- discrim_linear(mode = "classification") %>%
  set_engine("MASS")

lda_wf <- workflow() %>%
  add_model(lda_spec) %>%
  add_recipe(class_rec)

set.seed(123)
lda_cv <- fit_resamples(
  lda_wf, resamples = class_folds,
  control = control_resamples(save_pred = TRUE)
)
06

Experience

Full role-by-role detail lives in my resume, which I tailor per application. Here's the shape of it:

Aug 2025–Present
Data Analyst Graduate AssistantGVSU Charter Schools Office
Building the ETL & BI platform for 78 schools (see project above)
Mar 2025–Aug 2025
Study Abroad Peer AdvisorGrand Valley State University
Advised 100+ students; organized 17 years of program data
Dec 2023–Dec 2024
Associate Data EngineerCedar Gate Technologies
SQL/Redshift healthcare reporting under HIPAA
Aug 2023–Nov 2023
Data Analyst InternAyata Incorporation Pvt. Ltd
EDA & Power BI dashboards for stakeholder reporting
07

Education

DEC 2026

M.S. Data Science and Analytics

Grand Valley State University, Allendale, MI · GPA 3.9/4.0
AUG 2023

B.E. Computer Engineering

Tribhuvan University, Institute of Engineering, Kathmandu, Nepal
08

Let's talk

Open to full-time Data Analyst, Data Engineer, and Data Scientist roles and related roles, available starting December 2026. Open to relocation. Reach out any time.