What actually happens, in order, every frame.
This document answers: "when we optimize function X, what is the surrounding context?" For per-function detail see MASTER_FUNCTION_REFERENCE.md. For SH2 3D engine internals see sh2-analysis/SH2_3D_FUNCTION_REFERENCE.md.
Three CPUs run concurrently. Communication is exclusively through COMM registers
($A15120–$A1512E from 68K, $20004020–$2000402E from SH2).
┌─────────────────────────────────────────────────────────────────┐
│ 68K (7.67 MHz) Master SH2 (23 MHz) Slave SH2 (23 MHz)
│ 100% utilization 0–36% utilization 78% utilization
│ ← BOTTLENECK → (blocked waiting for (mostly async cmd
│ 68K commands) $27 pixel work)
│
│ V-INT fires every 16.67 ms (60 Hz, NTSC)
│ NTSC: 262 total scan lines (224 active + 38 retrace), 16.67 ms per frame
└─────────────────────────────────────────────────────────────────┘
COMM registers ($A15120–$A1512E on 68K / $20004020–$2000402E on SH2):
COMM0 HI ($A15120) Master SH2 dispatch index / "command in flight" flag
COMM0 LO ($A15121) Jump table index byte (written with COMM0 HI together)
COMM1 HI ($A15122) Slave work command (original game protocol)
COMM1 LO ($A15123) "Command done" signal managed by hw_init_short
COMM2:3 ($A15124) Source pointer / width parameter
COMM4:5 ($A15128) Destination pointer
COMM6 ($A1512C) Height + width-in-words packed
COMM7 ($A1512E) Slave SH2 doorbell ($0027 = cmd_27 work pending)
Fires at the start of every vertical blank (~60 Hz). Total overhead: ~300 cycles (excluding state handler). V-blank period: ~4,500 cycles. The 68K spends 49.4% of all cycles waiting for this interrupt to fire — see VBLANK_PERFORMANCE_ANALYSIS.md.
V-INT fires (V-blank period: ~4,500 cycles available)
│
├─ TST.W $FFC87A ; Any pending state work?
├─ BEQ → RTE ; No → exit immediately (fast path, ~12 cycles)
│
├─ SR = $2700 ; Disable all interrupts
├─ MOVEM.L → stack ; Save D0–D7 / A0–A6 (~120 cycles, 15 regs)
│
├─ D0 = [$FFC87A] ; Load dispatch state (pre-multiplied ×4)
├─ [$FFC87A] = 0 ; Clear — wait_for_vblank spin exits here
│
├─ A1 = jmp_table[D0] ; Load handler from table at ROM $0016B2
├─ JSR (A1) ; Call state handler (~24–1,600 cycles, see table)
│
├─ [$FFC964] += 1 ; Increment global frame counter
│
├─ MOVEM.L ← stack ; Restore all registers (~120 cycles)
├─ SR = $2300 ; Re-enable interrupts (H-INT allowed)
└─ RTE
The main loop sets $FFC87A before each V-INT to request a specific VDP/FB operation.
$FFC87A value |
Handler | Purpose |
|---|---|---|
| 0, 4, 8, 32 | vint_state_common ($001A64) |
VDP sync + Work RAM ops |
| 16 | vint_state_minimal ($001A6E) |
Quick VDP status read |
| 20 | vint_state_vdp_sync ($001A72) |
Full VDP register sync |
| 24 | vint_state_fb_toggle ($001C66) |
Frame buffer page flip |
| 28 | vint_state_sprite_cfg ($001ACA) |
VDP sprite configuration |
| 36 | vint_state_fb_setup ($001E42) |
Frame buffer pointer setup |
| 40 | vint_state_vdp_config ($001B14) |
VDP mode config |
| 44 | vint_state_transition ($001A64) |
Queue next state |
| 48 | vint_state_complex ($001BA8) |
Complex VDP operations |
| 52 | vint_state_fb_palette ($001E94) |
FB + palette update |
| 56 | vint_state_fb_dma ($001F4A) |
Frame buffer DMA |
| 60 | vint_state_cleanup ($002010) |
Clear SH2 command flags |
After the V-INT returns, the 68K runs the main game loop for the ~16 ms active display period. The loop has two phases: game logic and render submission.
Main loop (polling_loop / state_machine):
│
├─ poll_controllers ($00179E)
│ ├─ Read IO_DATA1 ($A10003), IO_DATA2 ($A10005)
│ ├─ Button remap via table at $FFEF82
│ └─ Store to $FFC86C (controller flags, longword)
│
├─ Game state dispatch via $FFC87E
│ ├─ State 0: Boot/menu (game_logic_entry $006200)
│ ├─ State N: Race mode (physics, AI, collision)
│ ├─ State M: Results / attract / name entry
│ └─ (each state calls its own sub-state machines)
│
├─ Object system update
│ ├─ obj_position_update ($007084) — integrate velocities
│ ├─ obj_velocity_x/y ($007F50/$007E7A) × 18 calls each
│ ├─ obj_collision_test ($007816) × 11 calls
│ └─ ai_entity_main_update_orch ($00A972) — AI steering/physics
│
├─ Render preparation
│ ├─ sh2_graphics_cmd ($00E22C) × 14 calls
│ │ └─ Builds block-copy command parameters from object list
│ │
│ └─ Render submission (see §3 below)
│
└─ Loop back to poll_controllers
| Address | Size | Name | Purpose |
|---|---|---|---|
$FFC87A |
word | vint_dispatch_state |
V-INT handler selector (pre-×4) |
$FFC87E |
word | game_state |
Main game state machine |
$FFC8A0 |
word | race_state |
Race sub-state |
$FFC964 |
long | frame_counter |
Global frame count (V-INT incremented) |
$FFC86C |
long | controller_flags |
Joypad button state (P1+P2) |
$FFC8AE |
word | effect_countdown |
Per-frame timer for effects |
$FFEF05 |
byte | boot_flag_1 |
Boot sequence gate |
$FFEF06 |
byte | boot_flag_2 |
Boot sequence gate |
$FFEF00 |
word | prng_seed |
Pseudo-random number generator state |
Every frame, the 68K submits commands to both SH2 CPUs. This is where ~60% of 68K cycles are spent — waiting in polling loops.
Each call submits one 2D block-copy command. The 68K is fully blocked until the Master SH2 finishes the copy.
68K side (sh2_send_cmd at $00E35A — B-004 v6-corrected):
│
├─ WAIT: tst.b COMM0_HI; bne.s .wait_ready ← WAIT #1: SH2 done?
│ (spins until Master SH2 clears COMM0_HI after previous copy + COMM cleanup)
│
├─ move.l A0, COMM3 COMM3:4 = source pointer
├─ move.l A1, COMM5 COMM5:6 = dest pointer
├─ move.b D1, COMM2_HI height (rows)
├─ move.b D0, COMM2_LO width/2 (word count)
├─ move.b #$22, COMM0_LO dispatch index (BEFORE trigger)
├─ move.b #$01, COMM0_HI trigger (LAST — Master SH2 sees this)
│ → dispatches to cmd22_single_shot at $023010F0
│
├─ WAIT: tst.b COMM0_LO; bne.s .wait_consumed ← WAIT #2: params read?
│ (Master SH2 clears COMM0_LO after reading COMM2–6)
│
└─ RTS (return — SH2 continues copy in background)
WAIT #1 of the NEXT call serves as the completion barrier.
2 COMM handshake waits (down from 3 in original protocol).
The SH2 clears COMM0_HI late (after copy + COMM1 cleanup), so WAIT #1 of
the next sh2_send_cmd implicitly waits for the previous copy to complete.
B-004 optimization (v6-corrected, DONE — verified 3600-frame autoplay): Reduced from 3 waits to 2 by having the SH2 clear COMM0_LO early (params consumed) and COMM0_HI late (copy complete). The 68K returns after WAIT #2, overlapping SH2 copy time with 68K game logic until the next sh2_send_cmd call.
These commands go to the Slave SH2 via COMM7 doorbell. The 68K does not wait for completion — the Slave processes async during the display period.
68K side (sh2_cmd_27):
│
├─ WAIT: tst.w $A1512E (COMM7); bne.s .wait ← only wait: Slave idle?
│ (spins until Slave clears COMM7 from previous cmd_27)
│
├─ move.l A0, $A15124 COMM2:3 = data pointer
├─ move.w D1, $A1512C COMM6_HI = height
├─ move.w D0, $A1512E COMM7 = $0027 (doorbell — triggers Slave)
│
└─ RTS (return immediately)
Slave side (inline_slave_drain @ SDRAM $06000608):
│
├─ Detect COMM7 != 0
├─ Copy COMM2:3, COMM6 to local registers
├─ Clear COMM7 = 0 (acknowledge)
└─ Execute pixel operation (add/OR/mask) on the region
→ Slave loops back to poll COMM7 for next doorbell
B-003 optimization (applied, working): Removed the synchronous 3-way handshake. 68K now fire-and-forgets, Slave processes all 21 calls overlapped with the 68K's next game logic pass.
Dispatch loop (runs continuously on Master SH2):
│
├─ R8 = $20004020 (COMM0 cache-through address)
│
├─ R0 = MOV.B @R8 ; read COMM0_HI
├─ CMP/EQ #0, R0 ; zero = no command?
├─ BT → loop ; yes → keep polling
│
└─ Non-zero:
├─ R0 = MOV.B @(1,R8) ; read COMM0_LO = command code
├─ SHLL2 R0 ; R0 *= 4 (jump table offset)
├─ R1 = $06000780 ; jump table base (SDRAM literal)
├─ R0 = MOV.L @(R0,R1) ; load handler address
├─ JSR @R0 ; call handler
└─ BRA → loop ; return to poll
Jump table at SDRAM $06000780 — key entries:
| COMM0_LO | Handler | Purpose |
|---|---|---|
$01 |
$060008A0 |
Standard 3-phase COMM6 handler (all original game cmds) |
$22 |
$023010F0 |
Expansion ROM: single-shot block copy (B-004) |
$27 |
COMM7 path | Routed to Slave via COMM7 (B-003), NOT dispatched via Master |
The Slave polls COMM1 (for original game scene commands) and COMM7 (for async pixel work). 78% utilization — Slave is busy for most of each frame.
Slave loop:
│
├─ Poll COMM1_HI ($20004022)
│ ├─ Non-zero → dispatch via jump table at $060005C8
│ └─ Zero → fall through to COMM7 check
│
├─ Poll COMM7 ($2000402E) [inline_slave_drain @ $06000608]
│ ├─ Non-zero ($0027) → process cmd_27 pixel work (async)
│ │ ├─ Read COMM2:3 (data ptr), COMM6 (dims)
│ │ ├─ Clear COMM7 = 0 (release 68K from its wait)
│ │ └─ Execute pixel region operation
│ └─ Zero → loop back to COMM1 poll
│
└─ Loop
NTSC: 16.67 ms/frame, 60 Hz target (actual: ~20–24 FPS at baseline).
| Time (ms) | CPU | Action |
|---|---|---|
| 0.00 | V-INT | Interrupt fires at $001684 |
| 0.00 | 68K | Check $FFC87A, dispatch VDP/FB state handler |
| 0.00 | 68K | Increment $FFC964 frame counter |
| 0.00 | 68K | Restore registers, RTE |
| 0.05 | 68K | poll_controllers — read joypad state |
| 0.05–12 | 68K | Game state dispatch (physics, AI, collision, menus) |
| 12–16 | 68K | Render preparation: sh2_graphics_cmd × 14 |
| 12–16 | 68K | Blocking: sh2_send_cmd × 14 (waits for Master SH2 each) |
| 12–16 | Master SH2 | Executes 14 block copies serially, each triggered by 68K |
| Throughout | Slave SH2 | Polls COMM7; processes cmd_27 × 21 async pixel ops |
| 16.67 | V-INT | Next interrupt fires; cycle repeats |
Why FPS is ~20–24 and not 60: The 68K blocks for Master SH2 completion on
every sh2_send_cmd call. The Master SH2's execution time (block copy) is added
to the 68K's frame time, not overlapped with it. Combined with 100% 68K
utilization for game logic, the frame cannot fit into 16.67 ms.
The master_sequencer ($00C200) drives scene transitions via $FFC87E:
master_sequencer ($00C200):
│
├─ Read game_state ($FFC87E)
├─ Dispatch via scene_state_dispatch ($00C662) → jump table
│
├─ Boot/attract: load assets, configure VDP, init objects
├─ Menu: input handling, track/car selection
├─ Race: full game loop (physics + SH2 render pipeline)
├─ Pause/options: freeze physics, show menu overlay
├─ Results: lap times, comparison, leaderboard
└─ Name entry: character input grid
Scene transitions → scene_transition ($00C870):
└─ Saves game_state, sets up next state, triggers V-INT state 11
(vint_state_transition) to queue next handler
During race mode, each frame the scene module calls:
sh2_frame_sync($00203A) — ensures SH2 is ready to receive new framesh2_framebuffer_prep($0027DA) — sets up 32X FB for new render- Game update + physics (all the game subcategories)
- Render command submission (14×
sh2_send_cmd+ 21×sh2_cmd_27)
Based on the above, here is where cycle savings can be found:
| Optimization | Cycles/Frame | Status |
|---|---|---|
B-003: sh2_cmd_27 fire-and-forget |
~3,000 saved | DONE |
B-004: sh2_send_cmd single-shot (v6-corrected) |
~1,400 saved | DONE (verified 3600-frame autoplay) |
B-005: Fire-and-forget sh2_send_cmd |
N/A | BLOCKED (COMM0_HI is frame sync barrier) |
| B-006: Slave SH2 parallel vertex transform | ~15–20% gain | REVERTED (COMM7 namespace collision) |
| B-009: FB write FIFO burst mode (2.4× raster) | TBD | OPEN |
Inline angle_to_sine (29 calls) |
~580 saved | Not started |
| Eliminate COMM0 busy-wait between commands | ~2,100 → 0 | Blocked by B-005 |
The architectural ceiling: the 3-way COMM handshake per sh2_send_cmd call
forces serialization of 68K and Master SH2. Removing these waits (via async
command queues or pipelining, i.e. B-005) is the highest-leverage single change.
See also:
- 68K_SH2_COMMUNICATION.md — COMM register protocol detail
- COMM_REGISTERS_HARDWARE_ANALYSIS.md — hardware hazards
- ARCHITECTURAL_BOTTLENECK_ANALYSIS.md — cycle budget
- MASTER_FUNCTION_REFERENCE.md — all 799 function entries