Skip to content

OPEN TO NEW OPPORTUNITIES

PORTFOLIO / 2026

ARLINGTON, VA / EASTERN TIME

SENIOR ASSOCIATE SOFTWARE ENGINEER / CAPITAL ONEFOUNDER / HIDDEN STRATA

Jose Alvarado AlvarengaI build systemsthat hold up.

Biotechnology-trained software engineer building reliable backend, cloud, and AI systems with Python, TypeScript, and AWS—shaped by work at Capital One and Merck, and by shipping my own products.

SPECIMEN 00CENTRAL DOGMA

STATIC MODEL

01

DNARAW ENCODING

02

mRNACODON / INSTRUCTION IR

03

PROTEINFOLDED FUNCTION

B.S. BIOTECHNOLOGY / MINOR IN COMPUTER SCIENCE

DRAG DNA / USE ← →

A software metaphor overlays the central dogma: DNA carries the ASCII binary for LIFE and CODE, messenger RNA expresses instruction mnemonics to load both values and call BUILD, and the folded protein resolves into the high-level expression BUILD LIFE, CODE. The biology labels are literal; the code abstraction is an analogy.
01EXPERIENCE
Capital One + MerckFINANCE / GMP

PROMOTER / Now

Currently in the lab.

A 12-week public series connecting the biotechnology degree with model internals, evaluation design, and reliable AI engineering.

ACTIVE BUILD — WEEK 3 NEXT

Protein language models, built from first principles.

Most model demos begin after the data is already split. This series starts one layer earlier: define a defensible homology-aware evaluation boundary, preserve negative results, and build each model from amino-acid bigrams toward a small masked Transformer.

Two tagged releases are public. Week 1 measured split contamination across 557,718 eligible proteins. Week 2 compared six unigram and bigram models under matched 100-million-pair training streams, with the sealed test left untouched.

CURRENT CHECKPOINT

Homology-aware data, checksummed manifests, and a sealed test firewall

Unigram, count-based bigram, and neural bigram baselines

Context MLPs and learned amino-acid embeddings next

A small masked Transformer and ESMC-300M comparison by Week 12

EXON 01 / Selected work

Work I can defend.

Products, infrastructure, and experiments with explicit contracts, measured outcomes, failure analysis, and reproducible evidence.

01LIVE

RepNotes AI

Founder · Hidden Strata LLC

Jan 2026 — Present

An AI-native iOS workout tracker that turns free-form workout notes into structured logs, progress analytics, insights, and next-workout plans.

PROBLEM

You're between sets with 60 seconds on the clock. Every workout app wants you to tap through menus and dropdowns to log what you just did. Notes apps are fast, but they leave you with a pile of unstructured text that can't track progress.

APPROACH

Log your workout the way you'd text a friend. Parser v2, an AI-assisted understanding pipeline, normalizes messy natural-language notes into exercises, sets, reps, weights, and cardio — and Plan AI turns that history into grounded insights and next-workout plans.

SELECTED HIGHLIGHTS

03 / 05

01

Parser v2: normalizes messy natural-language notes into exercises, sets, reps, weights, cardio, and training facts — with correction replay

02

Plan AI: facts-grounded coaching built on LLM planner/narrator stages, strict JSON contracts, and deterministic evidence validators

03

Rejects unsupported, stale, or fabricated recommendations before they ever reach the user

STACK

  • Swift
  • SwiftUI
  • OpenAI
  • Structured Outputs
  • Supabase
  • PostgreSQL
  • TypeScript
  • RevenueCat
  • Sentry
02IN PROGRESS

An experiment-driven series building protein language models from amino-acid bigrams toward a small masked Transformer, with homology-aware evaluation and reproducible local training.

PROBLEM

Random protein splits can place related sequences on both sides of an evaluation and inflate apparent generalization. Before deeper models are worth training, the dataset boundary, baselines, and test firewall have to be defensible.

APPROACH

Start with 557,718 eligible Swiss-Prot proteins, compare immutable random and UniRef50-group assignments under a pinned MMseqs2 audit, then build progressively from matched unigram and bigram baselines while preserving checksummed evidence.

SELECTED HIGHLIGHTS

03 / 05

01

Week 1: UniRef50 grouping lowered detected strong validation overlap from 87.4% to 35.9%, while still failing the frozen readiness gates

02

Published the negative result and prohibited both diagnostic assignments from model training

03

Week 2: evaluated unigram, count-based bigram, and neural bigram models across two matched 100M-pair training streams

STACK

  • Python
  • PyTorch
  • NumPy
  • UniProt
  • UniRef50
  • MMseqs2
  • ESMC-300M
  • uv
  • pytest
03OPEN SOURCE

WHOOP Personal MCP

Creator

2026 — Present

A provider-neutral, single-user MCP server that gives authorized AI clients read-only access to structured WHOOP wellness and training context.

PROBLEM

Biometric context is sensitive, time-dependent, and often incomplete. Connecting it to an AI client creates a larger trust boundary than a typical API wrapper: authorization, client identity, stale or missing values, and provider disclosure all need explicit behavior.

APPROACH

The difficult part was not fetching biometric data — it was creating a defensible trust boundary for sensitive information. The server combines owner-gated OAuth and consent, encrypted WHOOP tokens, strict redirect, Origin, and Host validation, explicit freshness and missing-data semantics, and a provider-neutral MCP interface without persisting API responses or tool results.

SELECTED HIGHLIGHTS

03 / 04

01

Native MCP 2026-07-28 Streamable HTTP with a stateless compatibility fallback for legacy clients

02

Six focused read-only tools plus optional event context, with explicit freshness, coverage, and missing-data behavior

03

PKCE, CIMD/DCR support, encrypted tokens, and strict redirect, Origin, and Host validation

STACK

  • TypeScript
  • Node.js
  • Express
  • SQLite
  • OAuth / PKCE
  • MCP 2026-07-28
  • Docker
  • GitHub Actions

EXON 02 / Experience

Reliable systems. Measurable outcomes.

Jan 2024 — Present

Senior Associate Software Engineer

Capital One · McLean, VA

Established governance and onboarding for an internal AI agent skills marketplace, defining initial requirements with the technical lead and onboarding 6+ engineering teams through hands-on PR reviews.

Translated engineering-team feedback into installation and skill-sync enhancements, then presented a GitHub metrics skill to 110+ engineers across Card Tech's AI Trailblazers program.

Architected an event-driven rewards pipeline with AWS CDK, Lambda, SQS, and SNS, replacing a two-week manual issuance process with one-day fulfillment and saving approximately 60 operational hours per month.

Delivered the first real-time customer and account enrichment for a travel messaging pipeline, enabling personalized send-time messaging for 100K+ customers across three marketing intents.

Co-led technical delivery and AWS Glue ETL optimization for multi-channel campaigns while maintaining zero production incidents.

Reduced distributed PySpark CI build times by 60–70%, from 15–20 minutes to approximately six, with pytest-xdist and shared session fixtures across 380+ tests.

Mentored 10 associate engineers across two cohort cycles on onboarding, code quality, and cloud architecture best practices.

Aug 2022 — Jan 2024

Associate Software Engineer

Capital One · Richmond, VA

Decoupled legacy decisioning services from Amazon RDS by migrating data access to DynamoDB-backed REST APIs, remediating security vulnerabilities and retiring infrastructure that cost $36K annually.

Led the modernization of three legacy APIs from Java 8 to Java 17, balancing technical debt with delivery priorities to reduce latency by 30% and improve stability.

Jun 2020 — Dec 2021

Associate Automation Engineer Co-op

Merck & Co. · Elkton, VA

Developed DeltaV distributed control system automation for Gardasil seed and production fermentation, plus PI ProcessBook and DeltaV Live operator graphics used under GMP documentation controls.

EDUCATION

B.S. Biotechnology, Minor in Computer Science

James Madison University · May 2022

CERTIFICATION

AWS Certified Solutions Architect — Associate

Amazon Web Services · Sep 2023

EXON 03 / About

Models write prose. Systems own the facts.

I build reliable production systems where correctness is not optional — across financial services at Capital One, GMP-regulated pharma at Merck, and an AI-native product through Hidden Strata.

My backend and cloud work centers on Python, TypeScript, and AWS. I also help engineering teams adopt internal AI agent tooling, from marketplace governance and onboarding across 6+ teams to skill-sync enhancements and a GitHub metrics skill presented to 110+ engineers.

For RepNotes AI, that same discipline becomes strict JSON contracts, deterministic validators, replayable evals, and production observability. In the protein-language-model series, it becomes immutable manifests, homology-aware evaluation, and a sealed test boundary. The biotechnology training and GMP controls taught me to treat traceability and failure modes as product requirements.

Languages

Python

TypeScript

Swift

Java

SQL

AI Engineering

OpenAI Structured Outputs

Claude / Claude Code

MCP

LLM Evals

Agent Tooling

Cloud & AWS

AWS CDK

Lambda

SNS / SQS

DynamoDB

Glue

Fargate

Railway

Backend & Data

Event-Driven Systems

Microservices

ETL Pipelines

PySpark

PostgreSQL / Supabase

Splunk

Product & Delivery

SwiftUI

RevenueCat

Docker

Sentry

OAuth 2.0

INTRON / OFF THE CLOCK

Ironman 70.3

2027 — race TBD

Richmond Half Marathon

after that

I'm training with my own tools in the loop — WHOOP Personal MCP exposes read-only recovery, sleep, and workout context with freshness and missing data made explicit. The data informs my judgment; it doesn't prescribe workouts or medical decisions.

3′ UTR / Contact

Let's build something grounded.

Open to new opportunities and collaborations — or just talking shop about AI products, MCP, and endurance training.

CTG·ATC·TTC·GAG

© 2026 Jose Alvarado Alvarenga