Semester-long research project
The project is the center of the course: a carefully scoped research effort that could mature into a top-tier architecture, systems, ML systems, robotics, or EDA paper. Teams of 1-2 are formed by bidding on a curated portfolio of directions; publication is an aspiration, not a grading requirement.
Track A
Computing for AI
Profile, serve, schedule, map, accelerate, or make reliable an LLM, agentic, physical, or neuro-symbolic workload.
Track B
AI for Computing
Build and rigorously evaluate an agent for software optimization, compilers, GPU kernels, architecture DSE, RTL/EDA, or verification.
Project Candidates (students are welcome to propose their own projects)
Module 1: Computing for AI
- SwarmServe - workflow-aware scheduling for multi-agent LLM serving
- CacheCraft - program-structure-aware KV-cache retention and tiering
- MuxFlow - small-language-model orchestration under real serving load
- AgentShield - fault tolerance for multi-step agentic LLM pipelines
- ActSpec - deadline-aware speculative action generation for real-time VLAs
- VLA-Roof - an analytical performance model for VLA inference
- FleetBatch - deadline-aware multi-robot serving on shared GPUs
- SceneCache - consumer-aware caching for 3D scene representations
- PlanFuse - profiling and co-scheduling hybrid VLA + motion-planning stacks
- FaultLine - joint adversarial and soft-error robustness for VLA control
- SymKern - demystifying and optimizing neuro-symbolic kernels on GPUs
- SysTwo - systems characterization of System-2 (test-time-compute) inference
- ESP-Reason - a reasoning-accelerator tile on the ESP SoC platform (FPGA)
- HoloCIM - a compute-in-memory noise-budget study for vector-symbolic AI
- HotPath - the kernel-to-system conversion rate of component speedups
- WattGuard - energy-SLO co-control for edge AI inference
Module 2: AI for Computing
- KernelScope - does profiler feedback make GPU-kernel agents better, per token?
- MapSmith - LLM agents vs. classical search for accelerator mapping
- FairDSE - a budget-matched showdown: LLM agents vs. BO/RL/GA for architecture DSE
- Signals - which feedback modality buys the most design quality per token?
- Ouroboros - an agent co-designs an accelerator for its own workload
- ProveGen - RTL generation with verification-in-the-loop, scored by mutation testing
- Spec2Socket - an agentic flow from spec to HLS to FPGA-measured accelerator
- LACE-Next - workload-driven agentic RISC-V instruction extension
- HDLGround - structured retrieval and grounding for hardware-design agents
- FlowPilot - log-reading LLM agents vs. black-box autotuners on OpenROAD
- SysDoctor - a planted-bottleneck benchmark for AI performance diagnosis
- AutopilotServe - a guarded closed-loop agent that keeps a serving stack tuned
- AutoEmbody - agents that configure VLA deployments for closed-loop success
- RedFlag - reward-hacking forensics and hardened audits for design agents
A titles-only preview of the curated project portfolio. Full briefs with research questions, baselines, evaluation plans, and platforms will be released at the Week 2 research marketplace; directions may be added, merged, or refined before bidding opens.
Minimum research standard
- A falsifiable research question and a precise intended claim.
- At least two meaningful baselines, including a classical or non-agentic baseline where applicable.
- A mechanism, not only a correlation or leaderboard result.
- At least one ablation or controlled intervention tied to the central claim.
- Joint reporting of task quality or success and relevant systems metrics.
- A held-out workload, configuration, system scale, or design budget.
- Explicit compute, GPU, API, token, simulation, and wall-clock budgets.
- Failure-case analysis, including invalid actions and tool failures for agentic design systems.
- Reproducible code, environment instructions, configurations, data, and figure-generation scripts.
Milestones
- P0 Project proposal
- P1 Infrastructure & baselines
- P2 Prototype & pilot results
- Midterm Midterm presentation (in class)
- P3 Evaluation plan & initial results
- P4 Main results & ablations
- P5 Complete draft, artifact & poster; final poster session (in class)
- Final Final paper & artifact
All written deliverables are due 11:59 PM ET. Check-ins P1-P5 are pacing devices, graded on completeness. No course deadline falls on the Thanksgiving holiday.
Final submission
An eight- to ten-page conference-style paper (excluding references and appendices); a repository with pinned environment, one-command smoke test, and a documented reproduction path for one central result; machine-readable results with scripts regenerating principal figures; an experiment manifest covering seeds, configurations, models, machines, tool versions, and resource budgets; a response-to-feedback memo and individual contribution statements.