Early-Career AI Engineer | Experimental AI Systems, Model Evaluation & ML Infrastructure
B.S. AI Engineering student, Western Governors University — Expected September 2027
San Marcos, Texas · montoyaraul34@gmail.com
GitHub Projects · Hugging Face · Limen Neural
Building reproducible experimental AI systems for post-training, synthetic data, model evaluation, and agentic engineering on hardware I can own.
I am an early-career AI engineer building experimental agentic systems, synthetic-data pipelines, model-evaluation tooling, post-training infrastructure, and model tools. I work in Python, Rust, CUDA, and related systems languages, with an emphasis on executable validation, reproducibility, provenance, and honest separation between shipped evidence and active experiments.
My current program connects three systems:
- Synthetic Factory evaluates and curates synthetic-data campaigns behind provenance, rights, and project-policy gates.
- Operation Prometheus reconstructs real software-engineering trajectories from issues, patches, reviews, fixes, CI, and merges.
- Agoge Model Forge turns eligible data into frozen training/evaluation contracts, local post-training runs, and held-out comparisons.
Frontier hosted-model outputs remain evaluation-only comparators in this program rather than weight-update data. Training candidates must independently pass the applicable provenance, license, and project-policy gates.
- UsagePal PR #56 — Antigravity metrics and Linux support — merged upstream. Added Linux Antigravity discovery and authentication paths, host-aware language-server metadata, deterministic process/socket discovery, quota-pool reporting, burn-rate status, and macOS-target dependency isolation.
- UsagePal PR #48 — Devin Linux support and richer telemetry — merged and shipped in the v0.7.70 release line. Added Linux credential discovery, Devin/Devin-Next authentication paths, ACU/quota/credit tracking, pace indicators, documentation, and tests.
- agent-afk PR #1019 — first-class xAI/Grok provider — merged. Added API-key and SuperGrok OAuth paths, provider selection, credential handling, documentation, and tests.
- agent-afk PR #1242 — xAI OAuth CLI compatibility — merged. Fixed proxy-version negotiation with validated overrides and fallbacks.
These contributions demonstrate work inside other maintainers' architectures: understanding existing contracts, passing review, and shipping changes upstream.
- Synthetic Data Factory — SpaceXAI Grok 4.6 — the Grok 4.6 experimental campaign, grouped as a collection of synthetic-data, provenance, and evaluation artifacts.
- Synthetic Data Factory — Anthropic Claude Fable 5 — the matched Claude Fable 5 experimental campaign and its public synthetic-data artifacts.
- Spikenaut — the neuromorphic program's models, datasets, telemetry, and event-language experiment artifacts.
- Grok-1 Dissection and Ternary Packing — checkpoint dissection, GOZ1 packing, quantization, and compression experiments organized in one collection.
The hosted frontier-model collections are evaluation artifacts; inclusion in a public collection does not make their contents eligible for model-weight updates.
The first controlled study will compare one open-weight starting checkpoint under matched conditions:
| Arm | Condition |
|---|---|
| A | Untouched starting checkpoint |
| B | Minimally curated, rights-cleared synthetic supervision |
| C | Synthetic-Factory-curated, rights-cleared synthetic supervision |
| D | Owned or permissively licensed real engineering trajectories |
| E | Curated synthetic and real-trajectory mixture |
The first run is local-first on an RTX 5080 16 GiB. The experiment contract will freeze model and tokenizer revisions, splits, seeds, metrics, leakage checks, and success/null criteria before results are inspected.
Current status: infrastructure and data contracts are being hardened; no training improvement is claimed in advance. Active work includes the Synthetic Factory rights-policy foundation, Agoge immutable readiness contracts, and the Operation Prometheus eligibility and quality ledger.
| Project | Evidence and purpose |
|---|---|
| synthetic-factory | Synthetic-data experimentation with append-only evidence, provider/channel provenance, strict validation, failure-oriented fixtures, evaluation-only frontier lanes, and active rights/project-policy enforcement |
| operation-prometheus | Event-sourced software-engineering trajectory forge spanning issue, implementation, review, repair, CI, and merge history |
| agoge-forger | Local-first PyTorch post-training and evaluation tooling with configuration, artifact, checkpoint, split, and readiness contracts |
| xai-dissect | Rust CLI for read-only structural analysis of open Grok-1 checkpoints, including tensor inventory and MoE routing maps |
| grok-ozempic | Grok-1 compression experiments centered on expert precision, routing fidelity, and residual drift |
| corinth-canal | Rust experimental pipeline connecting telemetry encoding, spiking state, projection, model inspection, MoE routing, and SAAQ validation |
| Spikenaut-SNN | From-scratch SNN runtime and export target for telemetry-driven neuromorphic experiments |
I also maintain modular Rust and neuromorphic libraries under Limen Neural, including neuromod, axon-encoder, nir-rs, synaptic-mesh, and brainstem-daemon.
| Area | Tools and methods |
|---|---|
| Languages | Rust, Python, Julia, CUDA C/C++ |
| ML systems | PyTorch, Hugging Face, LoRA/QLoRA, MoE, GGUF, Safetensors, quantization, SNNs |
| Data and evaluation | Immutable manifests, provenance and rights gates, held-out splits, leakage guards, deterministic and executable scoring, JSONL/Parquet artifacts |
| Infrastructure | GitHub Actions, Docker, Linux, self-hosted GPU CI, cloud-training scaffolding, automated review gates |
| Hardware | NVIDIA RTX 5080 16 GiB, AMD Ryzen 9 9950X, Fedora Linux |
My neuromorphic work is a specialist experimental track: event-driven state, telemetry-aware control, SNN runtimes, and FPGA-oriented export and validation. It complements the central agentic/model-evaluation portfolio without being required to understand it.
I use coding agents heavily, with isolated branches or worktrees, local tests, linting, benchmarks or experiment gates, pull-request review, repair, and human merge decisions. Agent assistance is attributed in commits, PRs, experiment notes, or release provenance where appropriate.
What I optimize for:
Measurable results · reproducibility · provenance · rights-aware lineage · fail-closed validation · leakage-resistant evaluation · explicit scope boundaries
- Land the end-to-end rights/project-policy boundary for Synthetic Factory and keep research-only outputs outside all weight-update paths.
- Freeze and execute the first local base-to-SFT held-out comparison using independently eligible data.
- Publish the resulting checkpoints, telemetry, evaluations, regressions, and null results with reproducible artifacts.
Primary author and maintainer: Raul Montoya Cardenas (rmems).
Recruiter-focused structure and experimental-systems portfolio edits were developed with OpenAI ChatGPT / Codex. Project-specific AI contributions remain attributed in their respective commits, pull requests, experiment records, and release provenance.



