Semester-long research project

The project is the center of the course: a carefully scoped research effort that could mature into a top-tier architecture, systems, ML systems, robotics, or EDA paper. Teams of 1-2 are formed by bidding on a curated portfolio of directions; publication is an aspiration, not a grading requirement.

Track A

Computing for AI

Profile, serve, schedule, map, accelerate, or make reliable an LLM, agentic, physical, or neuro-symbolic workload.

Track B

AI for Computing

Build and rigorously evaluate an agent for software optimization, compilers, GPU kernels, architecture DSE, RTL/EDA, or verification.

Project Candidates (students are welcome to propose their own projects)

Module 1: Computing for AI

  • SwarmServe - workflow-aware scheduling for multi-agent LLM serving
  • CacheCraft - program-structure-aware KV-cache retention and tiering
  • MuxFlow - small-language-model orchestration under real serving load
  • AgentShield - fault tolerance for multi-step agentic LLM pipelines
  • ActSpec - deadline-aware speculative action generation for real-time VLAs
  • VLA-Roof - an analytical performance model for VLA inference
  • FleetBatch - deadline-aware multi-robot serving on shared GPUs
  • SceneCache - consumer-aware caching for 3D scene representations
  • PlanFuse - profiling and co-scheduling hybrid VLA + motion-planning stacks
  • FaultLine - joint adversarial and soft-error robustness for VLA control
  • SymKern - demystifying and optimizing neuro-symbolic kernels on GPUs
  • SysTwo - systems characterization of System-2 (test-time-compute) inference
  • ESP-Reason - a reasoning-accelerator tile on the ESP SoC platform (FPGA)
  • HoloCIM - a compute-in-memory noise-budget study for vector-symbolic AI
  • HotPath - the kernel-to-system conversion rate of component speedups
  • WattGuard - energy-SLO co-control for edge AI inference

Module 2: AI for Computing

  • KernelScope - does profiler feedback make GPU-kernel agents better, per token?
  • MapSmith - LLM agents vs. classical search for accelerator mapping
  • FairDSE - a budget-matched showdown: LLM agents vs. BO/RL/GA for architecture DSE
  • Signals - which feedback modality buys the most design quality per token?
  • Ouroboros - an agent co-designs an accelerator for its own workload
  • ProveGen - RTL generation with verification-in-the-loop, scored by mutation testing
  • Spec2Socket - an agentic flow from spec to HLS to FPGA-measured accelerator
  • LACE-Next - workload-driven agentic RISC-V instruction extension
  • HDLGround - structured retrieval and grounding for hardware-design agents
  • FlowPilot - log-reading LLM agents vs. black-box autotuners on OpenROAD
  • SysDoctor - a planted-bottleneck benchmark for AI performance diagnosis
  • AutopilotServe - a guarded closed-loop agent that keeps a serving stack tuned
  • AutoEmbody - agents that configure VLA deployments for closed-loop success
  • RedFlag - reward-hacking forensics and hardened audits for design agents

A titles-only preview of the curated project portfolio. Full briefs with research questions, baselines, evaluation plans, and platforms will be released at the Week 2 research marketplace; directions may be added, merged, or refined before bidding opens.

Minimum research standard

  • A falsifiable research question and a precise intended claim.
  • At least two meaningful baselines, including a classical or non-agentic baseline where applicable.
  • A mechanism, not only a correlation or leaderboard result.
  • At least one ablation or controlled intervention tied to the central claim.
  • Joint reporting of task quality or success and relevant systems metrics.
  • A held-out workload, configuration, system scale, or design budget.
  • Explicit compute, GPU, API, token, simulation, and wall-clock budgets.
  • Failure-case analysis, including invalid actions and tool failures for agentic design systems.
  • Reproducible code, environment instructions, configurations, data, and figure-generation scripts.

Milestones

  1. P0 Project proposal
  2. P1 Infrastructure & baselines
  3. P2 Prototype & pilot results
  4. Midterm Midterm presentation (in class)
  5. P3 Evaluation plan & initial results
  6. P4 Main results & ablations
  7. P5 Complete draft, artifact & poster; final poster session (in class)
  8. Final Final paper & artifact

All written deliverables are due 11:59 PM ET. Check-ins P1-P5 are pacing devices, graded on completeness. No course deadline falls on the Thanksgiving holiday.

Final submission

An eight- to ten-page conference-style paper (excluding references and appendices); a repository with pinned environment, one-command smoke test, and a documented reproduction path for one central result; machine-readable results with scripts regenerating principal figures; an experiment manifest covering seeds, configurations, models, machines, tool versions, and resource budgets; a response-to-feedback memo and individual contribution statements.