Physics / Electronics & Microelectronics Digital Electronics I 100% Free Open Access
Chapter 7 • Theory & Derivations

Semiconductor Memory Architectures: SRAM, DRAM & Non-Volatile

Solid-state digital data storage architectures and device physics: memory organization, word length, bit capacity, 2D/3D matrix addressing, row/column address strobes (RAS/CAS); Static RAM (SRAM) 6-transistor (6T) CMOS bistable cell, precharge lines, read/write stability, and differential sense amplifiers; Dynamic RAM (DRAM) 1-transistor 1-capacitor (1T-1C) trench/stacked cell, destructive readout, capacitive leakage, and periodic refresh scheduling; Non-volatile memory technologies: Mask ROM, fuse PROM, UV-erasable EPROM, floating-gate tunneling EEPROM, and multi-level cell NAND/NOR Flash; SDRAM, DDR protocols, and multi-level cache memory hierarchy.

§7.1 Memory Organization, Matrix Addressing & Decoders

1. General Memory Architecture & Density Classification

A digital semiconductor memory stores binary data in a regular two-dimensional grid of binary memory cells. The overall storage capacity is expressed as:

$$\text{Capacity} = M \times N \quad (M \text{ words, each of } N \text{ bits})$$

To access an individual word among $M = 2^k$ addressable locations, the memory requires $k$ binary address input lines ($A_{k-1}, \dots, A_0$). For example, a $64\text{ K} \times 8$ memory chip possesses $2^{16} = 65,536$ words of 8 bits each, requiring $k = 16$ address lines, 8 bidirectional data lines ($D_7 - D_0$), and control lines: Chip Select ($\overline{CS}$), Output Enable ($\overline{OE}$), and Write Enable ($\overline{WE}$).

2. 2D Matrix Addressing & Row/Column Decoders

If $2^k$ words were laid out in a single linear column, the required address decoder would have $2^k$ output lines—requiring an astronomical number of gates ($65,536$ outputs for $k=16$). To achieve high physical density and compact square silicon layouts, memory arrays utilize 2D Matrix Coincident Addressing:

  1. The $k$ address bits are split into $r$ Row Address bits and $c$ Column Address bits ($k = r + c$).
  2. The Row Decoder activates exactly one horizontal Word-Line (WL) among $2^r$ rows, simultaneously enabling all $2^c$ storage cells along that row.
  3. The activated cells place their stored charges onto vertical Bit-Lines (BL).
  4. The Column Decoder controls a bank of column pass-gates (multiplexers) that route the selected bit-line data to the chip's output buffers.

A $64\text{ K}$-bit array arranged as a $256 \times 256$ matrix requires only one 8-to-256 row decoder and one 8-to-256 column decoder (512 total outputs instead of 65,536).

§7.2 Static RAM (SRAM): The 6T CMOS Memory Cell

1. Structure of the 6-Transistor (6T) CMOS SRAM Cell

Static RAM (SRAM) retains stored data indefinitely as long as DC power is maintained ($V_{DD} > 0$), without requiring periodic refresh cycles. The canonical 6T CMOS SRAM cell comprises:

  • Two cross-coupled CMOS inverters ($M_1, M_2, M_3, M_4$) forming a bistable latch with complementary internal storage nodes $Q$ and $\bar{Q}$.
  • Two nMOS access pass-transistors ($M_5, M_6$) connecting nodes $Q$ and $\bar{Q}$ to complementary bit-lines ($BL$ and $\overline{BL}$), gated by Word-Line ($WL$).

2. Operational Cycles of the 6T Cell

  1. Read Cycle:
    1. Precharge phase: Both $BL$ and $\overline{BL}$ are precharged to $V_{DD}$ (or $V_{DD}/2$) and then floated.
    2. Assertion phase: Word-line $WL$ is driven HIGH, turning on access transistors $M_5$ and $M_6$.
    3. Discharge phase: If $Q=0$ and $\bar{Q}=1$, node $Q$ discharges $BL$ through pull-down transistor $M_1$, creating a differential voltage swing $\Delta V = V_{BLB} - V_{BL} \approx 100 - 200\text{ mV}$.
    4. Sensing phase: A sensitive analog differential sense amplifier strobes, detecting $\Delta V$ and amplifying it rapidly to full CMOS logic levels ($0\text{ V}$ or $V_{DD}$).
  2. Write Cycle:
    1. Strong write drivers overdrive $BL$ and $\overline{BL}$ to opposite supply rails (e.g., $BL = 0\text{ V}, \overline{BL} = V_{DD}$ to write a 0).
    2. $WL$ is asserted HIGH. The strong pull-down on $BL$ overpowers the weaker internal pMOS pull-up transistor, flipping the cross-coupled latch into the new state.

Cell Sizing (Read Stability & Write Margin): To prevent the cell from accidentally flipping its state during a read operation (read disturbance), pull-down transistors $M_1, M_3$ must be made stronger than access transistors $M_5, M_6$ ($\beta_{\text{pull-down}} / \beta_{\text{access}} \ge 1.2 - 1.5$). Conversely, to ensure data can be written successfully, access transistors must be stronger than pull-up transistors $M_2, M_4$ ($\beta_{\text{access}} / \beta_{\text{pull-up}} \ge 1.0$).

§7.3 Dynamic RAM (DRAM): The 1T-1C Cell & Refresh Scheduling

1. The 1-Transistor 1-Capacitor (1T-1C) DRAM Cell

Robert Dennard (IBM, 1968) patented the 1T-1C DRAM cell, which slashed silicon area from 6 transistors down to a single access transistor $M$ and an integrated storage capacitor $C_s$ ($C_s \approx 25 - 35\text{ fF}$):

  • Logic '1' is stored as a packet of charge on $C_s$ ($V_C \approx V_{DD}$).
  • Logic '0' is stored as discharged state ($V_C \approx 0\text{ V}$).

Because the cell area is minuscule ($4F^2 - 6F^2$, where $F$ is the lithographic feature size), DRAM achieves gigabit storage densities orders of magnitude higher than SRAM, making it the universal choice for computer main system memory.

2. Charge Sharing & Destructive Readout

When Word-Line $WL$ is asserted during a read operation, the storage capacitor $C_s$ shares its charge with the much larger parasitic capacitance of the long bit-line $C_{BL}$ ($C_{BL} \sim 100 - 300\text{ fF} \approx 10 C_s$):

$$V_{\text{final}} = \frac{C_{BL} V_{\text{precharge}} + C_s V_C}{C_{BL} + C_s}$$

Precharging the bit-line to $V_{DD}/2$ produces a minute voltage perturbation:

$$\Delta V = \pm \frac{C_s}{C_{BL} + C_s} \left(\frac{V_{DD}}{2}\right) \approx \pm 100 - 150\text{ mV}$$

A cross-coupled regenerative sense amplifier detects $\Delta V$ and swings the bit-line fully to $V_{DD}$ or $0\text{ V}$. Because charge sharing partially discharges $C_s$, the readout is destructive. The sense amplifier must immediately rewrite the amplified logic level back into $C_s$ before closing the word-line (Restore Cycle).

3. Capacitor Leakage & Periodic Refresh Scheduling

Due to subthreshold MOSFET leakage and reverse-biased $p$-$n$ junction leakage currents, charge leaks off $C_s$ with a time constant of tens of milliseconds. To prevent catastrophic data loss, every row of the DRAM array must be read and rewritten (refreshed) periodically (typically every $64\text{ ms}$ at $85^\circ\text{C}$):

$$t_{\text{refresh interval}} \le 64\text{ ms}$$

Modern DRAM controllers interleave Distributed Auto-Refresh or Self-Refresh commands between normal read/write cycles, consuming less than $1 - 2\%$ of total memory bus bandwidth.

§7.4 Non-Volatile Memories: ROM, PROM, EPROM, EEPROM & Flash

1. Read-Only Memory (ROM) Evolution

Non-volatile semiconductor memories retain stored data indefinitely without requiring power supplies:

  1. Mask ROM: Data is permanently hardwired during wafer fabrication using a custom photolithographic contact mask. Highest density and lowest cost per bit in mass production, but zero programmability.
  2. Programmable ROM (PROM): Fabricated with microscopic Nichrome or polycrystalline silicon fuses in series with each memory cell. Programmed once by blowing selected fuses using high-current pulses (One-Time Programmable - OTP).
  3. Erasable PROM (UV-EPROM): Utilizes a Floating-Gate MOSFET (FGMOS) with a completely isolated conductive polysilicon gate embedded inside silicon dioxide dielectric. High-voltage pulses ($V_{PP} \approx 12 - 21\text{ V}$) inject electrons onto the floating gate via Hot-Carrier Injection (HCI). Stored electrons shift the transistor's threshold voltage ($V_t$), programming it to state 0. Erased by shining ultraviolet light ($254\text{ nm}$) through a quartz window on the chip package for 20 minutes, exciting electrons over the $\text{SiO}_2$ potential barrier.
  4. Electrically Erasable PROM (EEPROM): Employs ultra-thin tunnel oxide ($d_{\text{ox}} < 10\text{ nm}$) beneath the floating gate, enabling bidirectional electrical erasure via Fowler-Nordheim (F-N) Quantum Mechanical Tunneling. Can be erased and reprogrammed byte-by-byte in circuit.

2. Modern Flash Memory: NAND vs NOR Architectures

Invented by Fujio Masuoka (Toshiba, 1984), Flash memory erases blocks of cells simultaneously in a single flash operation via Fowler-Nordheim tunneling:

AttributeNOR FlashNAND Flash
Cell InterconnectionParallel (like NOR gate)Series strings of 32 to 128 cells (like NAND gate)
Random Access SpeedVery Fast ($50 - 80\text{ ns}$)Slow initial access ($25\text{ \mu s}$)
Serial ThroughputModerateExtremely High ($> 1\text{ GB/s}$)
Cell Area DensityLarge ($10F^2$)Extremely Compact ($4F^2$, 3D vertical stacked $> 200$ layers)
Primary ApplicationBIOS, Router Firmware (eXecute-In-Place)Solid-State Drives (SSDs), USB drives, Smartphones

§7.5 Advanced Memory Architectures: SDRAM, DDR & Cache Hierarchy

1. Synchronous DRAM (SDRAM) & Double Data Rate (DDR)

Early asynchronous DRAMs required address strobe handshakes ($\overline{RAS}, \overline{CAS}$) that bottlenecked high-speed microprocessors. Synchronous DRAM (SDRAM) synchronizes all control, address, and data lines to the master CPU system clock using internal pipelining and multi-bank prefetching.

Double Data Rate (DDR) SDRAM: Transfers data on both the rising and falling edges of each clock cycle, doubling throughput at identical clock frequencies:

  • DDR1: 2-bit prefetch buffer (2 data words per clock cycle).
  • DDR2: 4-bit prefetch buffer.
  • DDR3: 8-bit prefetch buffer.
  • DDR4: 16-bit prefetch buffer with independent memory bank groups.
  • DDR5: Dual 32-bit channels per module, on-die ECC, data transfer rates exceeding $6400\text{ MT/s}$.

2. Multi-Level Cache Memory Hierarchy

Because processor execution speed ($3 - 5\text{ GHz}$, sub-nanosecond cycles) is orders of magnitude faster than DRAM latency ($50 - 70\text{ ns}$), modern computer architecture incorporates a multi-tiered cache hierarchy exploiting the Principle of Locality (Temporal and Spatial):

$$\text{CPU Core Registers } (< 1\text{ ns}) \longrightarrow \text{L1 Cache (SRAM, } 32\text{ KB}, \sim 1\text{ ns}) \longrightarrow \text{L2 Cache (SRAM, } 512\text{ KB}, \sim 4\text{ ns}) \longrightarrow \text{L3 Cache (SRAM, } 32\text{ MB}, \sim 12\text{ ns}) \longrightarrow \text{Main Memory (DRAM, } 32\text{ GB}, \sim 60\text{ ns}) \longrightarrow \text{Secondary Storage (NVMe SSD)}$$

The Average Memory Access Time (AMAT) is:

$$\text{AMAT} = t_{\text{hit}} + \text{Miss Rate} \times \text{Miss Penalty}$$
Solved Problem Example 7.1: Design of a 64 KB Memory Subsystem Using 16 KB x 8 SRAM Chips

Design a $64\text{ KB}$ ($65,536 \text{ bytes}$) microprocessor memory subsystem using standard $16\text{ KB} \times 8$ SRAM memory chips. (a) Determine the number of memory chips required. (b) Calculate the total number of address bus lines required for the subsystem and how many address lines connect directly to each chip. (c) Design the chip-select address decoding circuit using a 2-to-4 binary decoder and specify the hexadecimal address range for each memory block.

Step 1: Calculate Chip Count
$$N_{\text{chips}} = \frac{\text{Total Capacity}}{\text{Chip Capacity}} = \frac{64\text{ KB}}{16\text{ KB}} = 4 \text{ memory chips}$$

Four 16 KB x 8 chips are required to synthesize 64 KB.

Step 2: Partition Address Bus Lines
$$64\text{ KB} = 2^{16} \implies 16 \text{ total address lines } (A_{15} - A_0). \quad 16\text{ KB} = 2^{14} \implies 14 \text{ address lines } (A_{13} - A_0) \text{ connect to each chip}$$

Address lines A0 through A13 connect in parallel to all 4 chips to select a byte within each chip.

Step 3: Design Decoder for Chip Select
$$\text{Remaining higher address lines } (A_{15}, A_{14}) \text{ connect to a 2-to-4 line active-LOW decoder to drive } \overline{CS}_0, \overline{CS}_1, \overline{CS}_2, \overline{CS}_3$$

A 2-to-4 decoder decodes A15 and A14 to enable one of the four chips.

Step 4: Determine Hexadecimal Memory Address Ranges
$$\begin{aligned} \text{Chip 0 } (A_{15}A_{14}=00): \ & 0000_{16} \text{ to } 3\text{FFF}_{16} \ (16,384 \text{ bytes}) \\ \text{Chip 1 } (A_{15}A_{14}=01): \ & 4000_{16} \text{ to } 7\text{FFF}_{16} \\ \text{Chip 2 } (A_{15}A_{14}=10): \ & 8000_{16} \text{ to } \text{BFFF}_{16} \\ \text{Chip 3 } (A_{15}A_{14}=11): \ & \text{C}000_{16} \text{ to } \text{FFFF}_{16} \end{aligned}$$

List continuous address space from 0000H to FFFFH.

Final Answer & Physical Insight

N_{\text{chips}} = 4; \quad 16 \text{ total address lines } (A_{13}-A_0 \text{ to chips}, \ A_{15}-A_{14} \text{ to 2:4 decoder}); \quad \text{Range: } 0000_{16} - \text{FFFF}_{16}

Solved Problem Example 7.2: DRAM Refresh Frequency, Time Overhead and Bandwidth Loss

An $8\text{ Gb}$ DDR4 DRAM memory chip is organized internally into $8192$ rows. All $8192$ rows must be refreshed within a maximum retention window of $t_{\text{ret}} = 64.0\text{ ms}$. Each individual row refresh cycle takes $t_{RC} = 45.0\text{ ns}$. (a) Calculate the average refresh frequency (how often a refresh command must be issued). (b) Calculate the total time consumed per $64\text{ ms}$ interval solely for refreshing. (c) Determine the percentage of memory bandwidth lost to refresh operations.

Step 1: Calculate Refresh Command Interval
$$t_{\text{REFI}} = \frac{t_{\text{ret}}}{N_{\text{rows}}} = \frac{64.0 \times 10^{-3}\text{ s}}{8192} \approx 7.8125 \times 10^{-6}\text{ s} = 7.81\text{ \mu s}$$

A refresh command must be executed every 7.81 microseconds.

Step 2: Calculate Total Time Spent Refreshing per 64 ms
$$t_{\text{refresh, total}} = N_{\text{rows}} \times t_{RC} = 8192 \times (45.0 \times 10^{-9}\text{ s}) \approx 3.6864 \times 10^{-4}\text{ s} = 368.64\text{ \mu s}$$

Multiply number of rows by row cycle time.

Step 3: Calculate Bandwidth Overhead Percentage
$$\text{Overhead} = \frac{t_{\text{refresh, total}}}{t_{\text{ret}}} \times 100\% = \frac{368.64\text{ \mu s}}{64,000\text{ \mu s}} \times 100\% \approx 0.576\%$$

Refresh consumes only 0.58% of total memory bandwidth.

Final Answer & Physical Insight

t_{\text{REFI}} = 7.81\text{ \mu s}, \quad t_{\text{total}} = 368.6\text{ \mu s per } 64\text{ ms}, \quad \text{Bandwidth Overhead} = 0.576\%

Solved Problem Example 7.3: Average Memory Access Time (AMAT) in Multi-Level Cache Hierarchy

A modern computer architecture features a two-level cache hierarchy: Level 1 ($L_1$) cache has an access time of $t_1 = 1.20\text{ ns}$ and a hit rate of $H_1 = 94.0\%$; Level 2 ($L_2$) cache has an access time of $t_2 = 6.00\text{ ns}$ and a local hit rate of $H_2 = 85.0\%$; Main DRAM memory has an access latency of $t_{MM} = 55.0\text{ ns}$. (a) Calculate the global miss rate for the cache system. (b) Calculate the Average Memory Access Time (AMAT) of the processor. (c) By what factor would memory performance degrade if the caches were eliminated?

Step 1: Calculate Global Miss Rate
$$\text{Miss Rate}_1 = 1 - 0.940 = 0.060. \quad \text{Global Miss Rate} = \text{Miss Rate}_1 \times (1 - H_2) = 0.060 \times (1 - 0.850) = 0.060 \times 0.150 = 0.0090 \ (0.90\%)$$

Only 0.9% of memory requests reach main DRAM.

Step 2: Calculate Average Memory Access Time (AMAT)
$$\text{AMAT} = t_1 + \text{Miss Rate}_1 \times [t_2 + (1 - H_2) \times t_{MM}] = 1.20 + 0.060 \times [6.00 + 0.150 \times 55.0]\text{ ns}$$

Formulate multi-level AMAT equation.

Step 3: Evaluate AMAT
$$\text{AMAT} = 1.20 + 0.060 \times [6.00 + 8.25] = 1.20 + 0.060 \times (14.25) = 1.20 + 0.855 = 2.055\text{ ns}$$

Average access time is only 2.06 ns.

Step 4: Compute Degradation Factor Without Caches
$$\text{Degradation} = \frac{t_{MM}}{\text{AMAT}} = \frac{55.0\text{ ns}}{2.055\text{ ns}} \approx 26.8 \times$$

Without caches, memory access would be nearly 27 times slower!

Final Answer & Physical Insight

\text{Global Miss Rate} = 0.90\%, \quad \text{AMAT} = 2.055\text{ ns}, \quad \text{Speedup over Raw DRAM: } 26.8\times

EXAM SUCCESS WORKSHOP

Solved University Examination Problems

Step-by-step mathematical solutions to classic university honors examination questions.