AI Systems Engineer / Florida Open to AI systems, evaluation, and research engineering roles
01

OpenAI × RunPod compute

Selected twice for OpenAI × RunPod compute credits during Parameter Golf, including a $1,000 grant for the final round.

02

OpenAI WebMCP Challenge

Three agent-native web apps submitted with 38 site-owned WebMCP tools between them.

03

Former Facebook contractor

Conversions API, Facebook Pixel, and custom JavaScript integrations.

Build · train · ship · operate

Model training.Agent-native web.Production systems.

I’m Francis Clase, an AI systems engineer. I train small language models under hard compute limits, build the tools agents use to act on the web, and run the production systems behind real businesses: email, calls, bookings, ads, and maps.

Portrait of Francis Clase

Francis Clase

Models · agents · production

FC / 2026

6

public Parameter Golf PRs

38

WebMCP tools shipped

5

GPU classes trained on

8,000+

search keywords grown

Model training / 01

OpenAI Parameter Golf.

Train the best language model that fits in a 16 MB artifact after ten minutes on eight H100s. The score is validation bits per byte on FineWeb. Lower is better.

OpenAI Parameter Golf project artwork
RunPod and OpenAI Parameter Golf compute credits banner

Compute credits · approved twice

  • March 20, 2026$25
  • April 30, 2026$1,000

Selected twice for OpenAI × RunPod compute credits during Parameter Golf, including a $1,000 grant for the final round.

Multiple rented systems in every GPU class used, from PCIe test boxes to eight-way SXM nodes, with compliance checked against the challenge rules before any compute was spent.

GPU classes

H100H200B200B300A100
Public pull requests · 6val_bpb · lower is better
  1. #309

    CLASE-Quant adaptive layer quantization

    1.1914

    Mar 21 2026

  2. #636

    11L XSA4 + EMA + GPTQ + FA3

    1.1234

    Mar 24 2026

  3. #1006

    JEPA + AdamW TTT + full GPTQ + FA3 + LZMA

    1.1085

    Mar 28 2026

  4. #1124

    v9 batched Muon + full GPTQ random calibration + JEPA research

    1.1194

    Mar 30 2026

  5. #2080

    BIJEPAX-lite JEPA + SP8192 CaseOps PPM

    0.9727

    Apr 30 2026

  6. #2083

    SP8192 CaseOps v13 PPM tuned gate, fresh 3-seed mean

    0.9418

    Apr 30 2026

Techniques across the six submissions

Adaptive layer quantizationFull-Hessian GPTQJEPA objectivesFlashAttention 3Batched MuonScore-first test-time trainingSentencePiece 8192 + CaseOpsLZMA / brotli artifact packingFused RMSNorm Triton kernels
View the public pull requests on GitHub

Agent-native web / 02

Three entries. 38 tools.

WebMCP is an emerging open standard that lets a website expose structured tools an AI agent can call directly, instead of guessing its way through the interface.

Challenge

The WebMCP Challenge, run by OpenAI

Window

August 25 – September 4, 2026

Field

7,058 registered participants · $35,000 prize pool

Supporters

Google Chrome · Cloudflare · Shopify · Vercel · Render · Netlify

The contract every entry publishes

document.modelContext.registerTool({
  name: "assess_route_weather",
  description: "Score the proposed route against live conditions",
  inputSchema: { /* closed JSON schema */ },
  execute: async (input) => { /* same engine the UI uses */ }
});
Aviation weather decision support15 tools

IMC Guardian

A second pair of eyes for a solo pilot. The decision stays human.

A pilot asks one plain-English question about a proposed route. Fifteen site-owned tools give the agent structured access to airport observations, forecasts, advisories, risk factors, alternates, configurable route-watch alerts, and an NTSB evidence replay. GPT-OSS 120B on Groq plans the tool calls. The browser executes them and shows every step. The app never issues a go or no-go.

React + ViteNOAA Aviation Weather CenterApple WeatherKitCesiumJS globeNASA GIBS satelliteGPT-OSS 120B on GroqVercel Functions
Two-sided booking marketplace15 tools

Booksy Reloaded

Set the rules once. Let the agent rebook inside them.

A bilingual beauty and wellness marketplace where the customer, the agent, and the provider share one visible appointment state. Fifteen tools rank providers with an inspectable score, apply a standing policy (usual provider, pay in person, stay under $50), and fail closed on any substitution. 25 of 25 benchmark runs completed through final booking.

Vanilla JS + ViteLeaflet + OpenStreetMapParameter lockingBrowser-local persistencePlaywright + VitestNative Chrome verification
Cleaning at conversation speed8 tools

Lander 5

A long intake form, rebuilt as eight tools an agent can call.

Built on the live quoting flow of a Florida cleaning company. The human page and the agent tools share one state engine: read policy, fill the request, calculate a quote, find a window, prepare a review, pause for approval, then reserve. Human edits invalidate stale quotes and approvals. Chrome 151 discovered and invoked all eight tools on the public page.

React + ViteShared state engineApproval gateBenchmark artifactGitHub ActionsNative Chrome verification

Developer tools & products / 03

Tools I build for myself, then ship.

An editor, an agent control plane, an API-to-everything generator, and my own map software. Built to remove friction from my own daily work first.

01 / Location intelligence / Field interfaceNative capture · 1814 × 900 · no upscaling
Custom location intelligence map covering a Florida metro region
Private build / 01

Flagship / Location intelligence

My own map software, built around decisions.

A location-intelligence product I built from scratch for Mac, Windows, and mobile workflows: custom search, place discovery, address context, local categories, territory views, and an interface designed for field decisions rather than sightseeing.

Geospatial UXApple MapsAddress dataCross-platformLocal search

Private product build · interface case study

Code editor · Code-OSS forkPublic

Shiriken IDE

A free editor built on VS Code’s open-source core with Claude Code, Codex, MiniMax, and NVIDIA-hosted models one click away. A Claude harness runs Groq, OpenRouter, Cerebras, and NVIDIA models inside the Claude Code interface. API keys live in the OS keychain. Quick Open launches any project in its own window, and an Agent Bridge schedules and coordinates Claude agents.

ElectronCode-OSSClaude CodeCodexNVIDIA NIMOpen VSX
Visit the Shiriken IDE site
Agent control planePrivate build

Claude Library

A private operations hub where AI agents talk to each other and to me. It holds a shared skills library, working memory, a command queue for local agents, cron schedules, a browser runner, and a spend ledger. A roll call pings Claude, GPT, and Gemini at once, and a three-tier escalation scanner reviews repositories with the cheapest model first.

Next.js 14MongoDBUpstash RedisPlaywright CDPPM2 agentTOTP auth

Private repository · available to discuss in detail

API contract → every interfacePublic

Fiber

A public developer-product concept that turns OpenAPI and JSON schemas into typed SDKs, live documentation, CLIs, MCP servers, and agent workflows, with reviewable diffs, governed releases, and evaluation traces.

OpenAPITyped SDKsMCP serversAgent toolsEval traces
Explore the live Fiber prototype

Production systems / 04

Real businesses run on these.

Systems I designed, built, and operate for live companies. Clients are described by industry rather than name. Every automation fails closed, logs who changed what, and has a manual override.

01

Meat market · Texas

Email platform that replaced Mailchimp

Designed and built front to back: a drag-and-drop block composer rendered server-side through MJML, subscriber segmentation, automations, templates, scheduling, and an AI writing assistant. Campaigns fan out through a message queue in batches with parallel sends over Amazon SES, live progress, retries, bounce and complaint handling, and suppression sync.

Next.jsMongoDBAmazon SESQStashMJMLBlock editor
02

Meat market · Texas

Call intelligence with live sentiment

Inbound calls are answered, transferred, and recorded through a programmable voice API, transcribed on a self-hosted Whisper server with a Groq fallback, then analyzed by Claude for sentiment, service quality, attention flags, and rep attribution. A dashboard polls live calls, tracks month-to-date revenue, ranks reps, and exports monthly PDF reports.

TelnyxWhisperClaude HaikuGroqQueue processingPDF reporting
03

Residential cleaning · Florida

Revenue operating system for a cleaning company

Hybrid quote and intake forms, a cents-based billing engine, fail-closed notification guards, audit logging, cleaner compensation, payroll, and admin dashboards. A Maps visibility loop scans local rankings on a schedule and adjusts Google Ads bids inside guardrails, with manual overrides on every automation.

Next.jsMongoDBGoogle Ads APIStripeOpenPhoneCron + diagnostics
04

Glass & auto service · Texas

Sales and revenue stack for a glass & auto service company

End-to-end intake, service logic, lead routing, scheduling, attribution, partner network, communications, and conversion-focused customer experiences, run from a password-protected admin command center with its own SDK and panels.

RevOpsCRMLead routingAttributionAdmin SDK
05

Field sales · Florida

Prospecting and lead management on a map

Geolocation-based lead organization with interactive maps, proximity calculations, status tagging, photo documentation, AI assessments, and mobile layouts that reorganize around the rep’s current position.

MapLibre GLGoogle PlacesGeolocationAI assessments
06

E-commerce · Texas

Merchant feed and ads automation

Live XML feed generation for Google Merchant Center with cost tracking, availability and shipping rules, image optimization, and bulk editing, paired with scripted Google Ads audits, budget and keyword changes, and overnight cleanup workflows.

Google Merchant APIVercel BlobGoogle Ads scriptsSharp

Agents & bots / 05

From Groq bots to custom OpenMausBot builds.

Agents that plan, tools that execute, and humans who approve. The same pattern runs my IDE, my ops hub, and the WebMCP entries.

01

Groq-hosted bots

Open-weight models such as GPT-OSS 120B and Nemotron on Groq, planning allowlisted tool calls that the browser or server executes.

02

Custom OpenMausBot builds

Personal builds of the open-source OpenMausBot chat app, where every contact is a real Claude or Codex agent with its own model, computer, and apps.

03

Claude agent harness

A harness that runs NVIDIA, Groq, OpenRouter, and Cerebras models inside the Claude Code interface, with preflight checks and live token counts.

04

Agent-to-agent chat

A shared chat hub where Claude, GPT, and Gemini agents post status, pick up tasks, and answer roll calls.

05

Tiered model escalation

Repository scans that start with the cheapest model, verify with a second, and send only confirmed findings to the strongest.

06

Custom AI video

Deterministic, seek-safe video compositions built with HyperFrames for product and demo reels.

Evaluation workflow

Task design and evaluation, grounded in real operating data.

01

Frame the work

Turn a real operating problem into a bounded task with source systems, constraints, and a defensible target answer.

02

Build the environment

Connect CRM, account data, email, chat, maps, runbooks, dashboards, or APIs into a realistic working context.

03

Define correctness

Write objective scoring criteria, edge cases, escalation rules, and failure conditions that can be checked, not hand-waved.

04

Calibrate the agent

Run, inspect, refine, and re-test until the task is difficult for the right reasons and reliable in production.

Task authoringAgent evaluationScoring rubricsRunbooksCRM dataSupport triage

Model operations / 06

10B+

Estimated cumulative tokens used

Fluent across frontier and open-weight models.

High-volume, hands-on use across research, coding, agent orchestration, evaluation, automation, debugging, and production delivery. Not occasional prompt testing.

01

OpenAI Codex Pro

02

ChatGPT Pro / GPT-5.6

03

Claude models

04

Gemini Pro

05

Open-weight LLMs

Capabilities / 07

Broad range. One operating standard.

Technical depth across the entire path from training and acquisition to automation, intelligence, and evaluation.

01

Model training

Small-model pretraining under hard artifact and time budgets, quantization, test-time training, custom Triton kernels, and multi-node runs on rented H100, H200, B200, B300, and A100 systems.

02

Agent-native web

WebMCP tool design, closed input schemas, approval gates, shared human-and-agent state, native browser verification, and benchmark artifacts.

03

AI systems

Custom Claude agents, agent-to-agent orchestration, agent harnesses, task environments, model evaluation, scoring rubrics, and human review.

04

Revenue operations

Salesforce-certified CRM architecture, lead routing, account reconciliation, fulfillment workflows, dashboards, scheduling, and lifecycle email.

05

Data & location

Geospatial interfaces, custom maps, address intelligence, territory analysis, local search, and data-quality operations.

06

Growth engineering

Meta Conversions API, Facebook Pixel, TikTok and paid-media automation, SEO systems, reputation management, and prospecting.

07

Product engineering

Python-certified development, Next.js 14–16, TypeScript, JavaScript, APIs, e-commerce automation, and iOS/Android products.

Contact / 08

Projects, roles, and collaboration.

Model training, agent evaluation, agent-native products, revenue operations, location intelligence, or a system that has to work across many services. I’m interested in concrete problems with measurable outcomes.

Francis Clase

AI systems, agent evaluation, location intelligence, revenue operations, product engineering, and independent machine-learning research.

© 2026 Francis ClaseBased in Florida · Shared publicly