Close
0%
0%

Building a GPU from scratch in Verilog

A complete graphics pipeline written in RTL Verilog, built by a single developer using open source tools. Not a soft-core. Not an emulator.

Similar projects worth following
Most people assume designing a GPU requires hundreds of engineers, sub-10nm fabs, and million-dollar budgets. This project challenges that assumption. NovaGPU TS 1T is a complete GPU in synthesizable Verilog RTL, designed from scratch with a custom token dataflow architecture called N.E.O.N. Instead of a scheduler dispatching instructions, the data itself triggers execution — no warp scheduler, no idle cycles. The pipeline includes: PCIe 4.0, Token Matching Unit, Shader Cluster with custom 8-opcode ISA, hardware BVH ray tracing (TTU), frame generation without AI (MVU), Pineda triangle rasterizer, dual-port SRAM, and specialized accelerators for vegetation, water, geometry prediction, and post-processing. Current state: 14 RTL modules complete, synthesis and place & route successful, targeting Artix-7, 47/48 tests passing, 3D animations produced directly from RTL simulation. One developer. Open source tools. MIT license. Repository: https://github.co

NovaGPU TS1T is an experimental graphics processing architecture written in Verilog RTL, designed for research, simulation, and validation on FPGAs, with potential future evolution to ASICs.

The project explores a combination of classic rasterization, token-based processing, specialized accelerators, and image generation directly from RTL simulation.

Verification Status
Latest results shown:

Total tests: 48
Tests passed: 47
Tests failed: 1
Success rate: 97%
Only one failure pending:

A5 — Handling degenerate triangles within the Triangle Rasterizer
Currently demonstrated capabilities
Rasterization
Functional triangle rasterizer.
Bounding-box traversal.
Pineda-like edge functions.
Write to the framebuffer.
Validated pixel generation.
Geometry
Visual demonstrations:

Rotating 3D cube.

Rotating 3D Tetrahedron.

Generated from RTL simulation.

Visual Generation
The current workflow is:

Repository ↓ RTL Compilation ↓ Simulation ↓ Framebuffer ↓ PPM Export ↓ MP4 Conversion ↓ Final Video

Not used:

Blender
Unity
Unreal Engine
The image comes from the framebuffer generated by the simulation.

Main Components
TMU
Token Matching Unit

Responsible for:

Token Reception
Tag Matching
Data Flow Coordination
Shader Cluster
Executes shading operations.

Tests:

NOP
ADD
MVP Transform
Currently validated.

Triangle Rasterizer
Fragment Generation.

Functions:

Triangle rasterization
Bounding boxes
Pixel coverage
Framebuffer writing
Currently the module closest to 100% implementation.

BVH
Bounding Volume Hierarchy

Validated for:

Traversal
Ray hit
Ray miss
SRAM Controller
Validated for:

Write
Read
Ack
Framebuffer integration
Budget Controller
Computational budget control for complex tasks.

MVU
Motion Vector Unit

Designed for:

Reuse of temporal information
Possible frame interpolation
Motion optimization
New Modules
AFA
Aquatic & Foliage Accelerator

Specialized in:

Vegetation
Procedural motion
Waves
Environmental elements
Testing:

6/6
MPE
Meta Prediction Engine

Experimental prediction system.

Capabilities:

Observation History
Predictions
Prefetch
Tests:

5/5
GIA
Geometry Intelligence Accelerator

Experimental module for:

Geometric prediction
Object tracking
Spatial pattern detection
Tests:

4/4
Nexus
Recently introduced.

Based on what has been shown so far, it appears to be geared towards coordination and interconnection between internal blocks.

PVA
Another new block recently incorporated.

Detailed public documentation is still lacking.

FPGA Synthesis
Previously validated on:

Xilinx Artix-7
Results shown:

Successful synthesis
Successful implementation
Bitstream generated
Next Goal
Based on what you have mentioned several times:

Physical FPGA
Immediate goal:

Tang Nano 9K
Not for running a full GPU.

But to verify:

RTL Correctness
Real Integration
Timing
FPGA → ASIC Flow
Current Demonstration Goal
Generate a complete sequence:

10 seconds
60 FPS final video
Approximately 600 frames
Starting from:

Fast PPM generation
Automatic export
MP4 conversion
With more complex scenes than the initial demonstrations.

Project Philosophy
What distinguishes NovaGPU TS1T from many educational projects is that it attempts to build a complete graphics architecture using open and reproducible RTL:

Source code available.
Reproducible simulation.
Validated synthesis.
Demonstrable visual generation.
Planned path towards FPGA and eventually ASIC.

Currently, the strongest indicator of progress is that the validation went from approximately 33/37 tests (~89%) to 47/48 tests (~97%), with only one known active fault remaining in the rasterizer.

  • 47/48 TEST

    Jostin06/11/2026 at 14:32 0 comments

    47/48 Tests Passing - Only One Bug Left :p

    Today I finally reached 47 out of 48 tests passing on NovaGPU TS1T.

    Honestly, I didn't expect to get this far this quickly.

    Over the last few days I added several new modules and fixed a lot of integration bugs across the RTL.

    New modules currently implemented:

    • Nexus
    • PVA
    • MPE
    • AFA
    • GIA

    Current validation results:

    • Total tests: 48
    • Passed: 47
    • Failed: 1
    • Success rate: 97%

    The only remaining failure is:

    • A5 – Degenerate Triangle Handling

    Ironically, it could be the easiest bug in the entire project or the hardest one. I honestly have no idea yet XD.

    The project is still focused on:

    • FPGA validation
    • Future ASIC exploration
    • RTL-generated graphics
    • Reproducible simulation flow

    One of the things I like most is that the current visual demos are generated directly from RTL simulation.

    No Blender. No Unity. No external graphics engine.

    Current demos:

    • Rotating 3D cube
    • Rotating tetrahedron
    • Automatic PPM generation
    • MP4 video export

    The complete flow currently looks like this:

    Clone repository ↓ Compile RTL ↓ Run simulation ↓ Generate framebuffer ↓ Export PPM frames ↓ Convert to MP4

    My next goal is improving frame generation speed and producing longer animations directly from the RTL pipeline.

    I also plan to scale portions of the design for testing on a Tang Nano 9K FPGA.

    Repository:

    https://github.com/nova-studios-hw/novagpu-ts1t

    Discord:

    https://discord.gg/RfQwz8ySr

    Thanks for reading

View project log

Enjoy this project?

Share

Discussions

Similar Projects

Does this project spark your interest?

Become a member to follow this project and never miss any updates