An arithmetic logic unit (ALU) is the block inside a processor that performs arithmetic and logic operations on binary data, selected by an operation code. Writing one in Verilog is a standard exercise because it combines a case statement, arithmetic operators, flags and a clean interface, and it is small enough to verify completely. Below is a synthesizable 8-bit ALU with a self-checking testbench.
Specification
- Two 8-bit operands
aandb, a 4-bit operation codeop. - Operations: add, subtract, AND, OR, XOR, NOT, shift left, shift right, increment, decrement, compare.
- Outputs: 8-bit
resultand flags: zero, carry, negative, overflow. - Purely combinational; a register stage can be added outside.
Operation codes
| op | Operation | Result |
|---|---|---|
| 0000 | ADD | a + b |
| 0001 | SUB | a − b |
| 0010 | AND | a & b |
| 0011 | OR | a | b |
| 0100 | XOR | a ^ b |
| 0101 | NOT | ~a |
| 0110 | SHL | a << 1 |
| 0111 | SHR | a >> 1 |
| 1000 | INC | a + 1 |
| 1001 | DEC | a − 1 |
| 1010 | CMP | flags from a − b, result = 0 |
Verilog code for the ALU
module alu #(parameter W = 8) (
input wire [W-1:0] a,
input wire [W-1:0] b,
input wire [3:0] op,
output reg [W-1:0] result,
output reg carry, // carry/borrow out of add/sub
output wire zero, // result == 0
output wire negative, // result MSB (signed)
output reg overflow // signed overflow on add/sub
);
localparam ADD = 4'b0000, SUB = 4'b0001, AND_ = 4'b0010, OR_ = 4'b0011,
XOR_ = 4'b0100, NOT_ = 4'b0101, SHL = 4'b0110, SHR = 4'b0111,
INC = 4'b1000, DEC = 4'b1001, CMP = 4'b1010;
reg [W:0] tmp; // one extra bit to capture carry
always @* begin
// defaults prevent latches
result = {W{1'b0}};
carry = 1'b0;
overflow = 1'b0;
tmp = {(W+1){1'b0}};
case (op)
ADD: begin
tmp = {1'b0, a} + {1'b0, b};
result = tmp[W-1:0];
carry = tmp[W];
overflow = (a[W-1] == b[W-1]) && (result[W-1] != a[W-1]);
end
SUB, CMP: begin
tmp = {1'b0, a} - {1'b0, b};
result = (op == CMP) ? {W{1'b0}} : tmp[W-1:0];
carry = tmp[W]; // 1 = borrow
overflow = (a[W-1] != b[W-1]) && (tmp[W-1] != a[W-1]);
end
AND_: result = a & b;
OR_: result = a | b;
XOR_: result = a ^ b;
NOT_: result = ~a;
SHL: {carry, result} = {a, 1'b0};
SHR: {result, carry} = {1'b0, a};
INC: {carry, result} = {1'b0, a} + 1'b1;
DEC: {carry, result} = {1'b0, a} - 1'b1;
default: result = {W{1'b0}};
endcase
end
assign zero = (op == CMP) ? (tmp[W-1:0] == {W{1'b0}}) : (result == {W{1'b0}});
assign negative = (op == CMP) ? tmp[W-1] : result[W-1];
endmoduleDesign notes
- Default assignments at the top of the
always @*block give every output a value on every path, so synthesis infers pure combinational logic and no latches. - The extra bit in
tmpcaptures the carry out of addition and the borrow out of subtraction without a separate comparator. - Signed overflow is detected from the operand and result sign bits: for addition, overflow occurs when both operands have the same sign and the result’s sign differs.
- CMP computes the subtraction only for the flags and forces the result to zero, which is how processors implement compare instructions.
- Concatenation (
{carry, result} = {a, 1'b0}) implements the shifts and captures the shifted-out bit in one line. - Parameter W makes the same code work for 16- or 32-bit ALUs.
Self-checking testbench
`timescale 1ns/1ps
module tb_alu;
localparam W = 8;
reg [W-1:0] a, b;
reg [3:0] op;
wire [W-1:0] result;
wire carry, zero, negative, overflow;
integer errors = 0, i;
alu #(.W(W)) dut (.a(a), .b(b), .op(op), .result(result),
.carry(carry), .zero(zero), .negative(negative),
.overflow(overflow));
task check(input [3:0] t_op, input [W-1:0] t_a, t_b,
input [W-1:0] exp_res, input exp_c);
begin
op = t_op; a = t_a; b = t_b; #1;
if (result !== exp_res || carry !== exp_c) begin
errors = errors + 1;
$display("FAIL op=%b a=%h b=%h got res=%h c=%b exp res=%h c=%b",
t_op, t_a, t_b, result, carry, exp_res, exp_c);
end
end
endtask
initial begin
// directed vectors
check(4'b0000, 8'h0F, 8'h01, 8'h10, 1'b0); // ADD
check(4'b0000, 8'hFF, 8'h01, 8'h00, 1'b1); // ADD with carry out
check(4'b0001, 8'h10, 8'h01, 8'h0F, 1'b0); // SUB
check(4'b0001, 8'h00, 8'h01, 8'hFF, 1'b1); // SUB with borrow
check(4'b0010, 8'hF0, 8'h3C, 8'h30, 1'b0); // AND
check(4'b0011, 8'hF0, 8'h3C, 8'hFC, 1'b0); // OR
check(4'b0100, 8'hF0, 8'h3C, 8'hCC, 1'b0); // XOR
check(4'b0101, 8'hAA, 8'h00, 8'h55, 1'b0); // NOT
check(4'b0110, 8'h81, 8'h00, 8'h02, 1'b1); // SHL, carry = old MSB
check(4'b0111, 8'h81, 8'h00, 8'h40, 1'b1); // SHR, carry = old LSB
check(4'b1000, 8'hFF, 8'h00, 8'h00, 1'b1); // INC wraps
check(4'b1001, 8'h00, 8'h00, 8'hFF, 1'b1); // DEC wraps
// random add/sub against a reference model
for (i = 0; i < 200; i = i + 1) begin
a = $random; b = $random;
check(4'b0000, a, b, a + b, ({1'b0,a} + {1'b0,b}) >> W);
check(4'b0001, a, b, a - b, (a < b));
end
if (errors == 0) $display("PASS: ALU tests passed");
else $display("FAIL: %0d errors", errors);
$finish;
end
endmoduleRun it on any Verilog simulator. The structure (a check task, directed vectors, then random vectors against a reference expression) is the pattern explained in the Verilog testbench tutorial.
Extensions
- Add multiply and divide (multi-cycle, with a
validoutput). - Register the inputs and outputs and measure the maximum clock frequency after synthesis.
- Replace the ripple adder inferred by
+with a carry-lookahead adder and compare timing. - Integrate the ALU into a simple RISC-V datapath; see the RISC-V design and verification course.
- Write the testbench in SystemVerilog with constrained-random operands and functional coverage.
Learn RTL design properly
Blocks like this ALU are the first project in our RTL design course, where they are synthesized, timed and reviewed to industry coding standards. For more practice, see Verilog interview questions and VLSI projects for ECE students.
Frequently asked questions
What is an ALU in Verilog?
A combinational module that performs arithmetic and logic operations on its operands according to an operation code, typically written with a case statement inside an always @* block.
Why use default assignments in the ALU’s always block?
So that every output is assigned on every path through the case statement. Without them, synthesis infers latches to hold outputs that are not assigned for some opcodes.
How is the carry flag generated?
By computing the operation one bit wider than the operands; the extra most-significant bit is the carry out of addition or the borrow out of subtraction.
How is signed overflow detected?
For addition, overflow occurs when both operands have the same sign and the result has the opposite sign; for subtraction, when the operands have different signs and the result’s sign differs from the first operand.
Is this ALU synthesizable?
Yes. It uses only combinational constructs, arithmetic and logic operators, and a fully assigned case statement.
Share your question in comments or talk to our mentor team for batch guidance.
Ask the Admin Team
Drop your basic question in comments: eligibility, prerequisites, tools, fee range, and placement support.
Our team reviews and responds regularly.
