RTOS Architecture: How to Design Tasks, Queues, Events, Timers and Synchronisation

RTOS Architecture: How to Design Tasks, Queues, Events, Timers and Synchronisation

Using an RTOS doesn’t make your architecture better. It moves the difficulty.

A system with twenty tasks, fifteen queues, eight mutexes, three event groups and a timer service can be considerably harder to reason about than a superloop, and the failure modes are worse — a superloop that misses a deadline misses it visibly, every cycle. An RTOS that misses one does it under a specific interleaving, once a week, on the unit with the most network traffic.

Article 02 handed you a task table. Nine tasks, nine priorities, nine stack sizes, and two rules stated flatly: priority follows deadline rather than importance, and a task that can block never sits above one that can’t.

I asserted those. This article derives them, using real scheduling analysis on the KIC-400 — and then covers the three places that analysis stops being true on hardware you can buy.


The Concurrency Model Comes First

The question is never “how many tasks.” It’s what independent threads of execution the product actually has, and that’s answerable from requirements before you open an editor.

Five questions define a task. If you can’t answer all five, you don’t have a task, you have a function that hasn’t found its home:

  1. What wakes it? A timer, a queue, an ISR notification, a socket.
  2. What’s its deadline? A number, or explicitly none.
  3. What can block it, and for how long? Bounded, always.
  4. What does it own? Resources, state, or nothing.
  5. What happens if it stops? Someone has to notice. Usually the health monitor.

Article 02’s table is those five columns. That’s all a task table is.

Five questions that define a task: what wakes it, its deadline, what can block it, what it owns, and what happens if it stops.
Five questions define a task.

Deriving the Task Set

Four reasons justify a separate task. Not “it’s a separate module” — that’s article 03’s task-per-function trap, and it produces forty tasks and 40 KB of stack.

A distinct deadline. Control runs at 1 kHz with a 50 µs jitter budget. Diagnostics runs at 10 Hz with no hard deadline. They cannot share a task, because the shared task inherits the tighter deadline and the looser one’s worst-case execution time.

A distinct blocking behaviour. MQTT blocks inside a TLS handshake for seconds. Modbus RTU must respond within milliseconds. Same task means the Modbus response waits for a handshake.

A distinct failure domain. REQ-06 said a sensor failure must not stop the control loop. Merge sensing into control and that requirement is unsatisfiable — not degraded, unsatisfiable.

A distinct owned resource. The SPI bus manager owns SPI2 and serialises access to it. That serialisation is the task.

Run those four over the KIC-400 requirements and the nine tasks fall out. Two boundary calls worth showing, because they’re the ones that go the other way:

Diagnostics and Health Monitor merged. Both periodic, both low priority, neither blocks, same failure domain. Two tasks would cost 768 bytes of stack and buy nothing. Merged.

Modbus RTU and Modbus TCP kept separate, despite implementing nearly the same register map. Different deadlines (5 ms vs 100 ms), different blocking behaviour (none vs sockets), different failure domains. Same protocol logic, called from two contexts — the code is shared, the execution context isn’t. That distinction is the one people miss when they say “don’t duplicate the Modbus stack.”

Diagnostics and Health Monitor merged into one task; Modbus RTU and Modbus TCP kept as separate tasks that call shared register-map code.
Shared code, separate execution contexts.

Priority Assignment, Properly

Here’s where article 02’s first rule gets its proof.

Rate monotonic, in one paragraph

For periodic tasks with deadlines equal to their periods, the optimal fixed-priority assignment is: shorter period gets higher priority. Not more important. Shorter period. If any fixed-priority assignment can meet all the deadlines, rate-monotonic can — that’s Liu and Layland, 1973, and it’s the whole basis for the rule.

The schedulability bound for n tasks:

U = Σ (Cᵢ / Tᵢ)  ≤  n · (2^(1/n) − 1)

For nine tasks that’s 0.72. Under 72% total utilisation, rate-monotonic priorities guarantee every deadline is met. Above it, you need the sharper test below — it might still be fine.

On the KIC-400

Budgets first, measurements after. This is what the table looks like before anyone writes code:

TaskPeriod TBudget CU = C/TRM priority
Control1 ms60 µs0.0606 (highest)
Modbus RTU5 ms (sporadic)300 µs0.0605
Sensor10 ms200 µs0.0205
lwIP tcpip_thread1 ms (modelled)100 µs0.1004
Modbus TCP50 ms800 µs0.0163
MQTT5 s40 ms0.0083
Health Monitor100 ms150 µs0.0022
Firmware Update———2
Logging———1
Total0.266

27% against a 72% bound. Comfortable, and the headroom is deliberate — on an M7 with caches you want real margin, for reasons in the next section.

Stacked utilisation bar for the KIC-400 tasks totalling 0.266, compared to a 50 percent practical ceiling and the 0.72 rate-monotonic bound for nine tasks.
27% utilisation against a 72% bound.

Notice Firmware Update at priority 2 with no period and no utilisation. It’s the most commercially important feature on the list. It has no deadline, so rate-monotonic gives it nothing. That’s article 02’s rule, and now it’s not an opinion.

When the utilisation bound isn’t enough

The bound is sufficient, not necessary — fail it and you might still be schedulable. The exact test is response-time analysis:

Rᵢ = Cᵢ + Bᵢ + Σ  ⌈Rᵢ / Tⱼ⌉ · Cⱼ        for all j of higher priority

Iterate until it converges. If Rᵢ ≤ Dᵢ, the task meets its deadline.

Bᵢ is the blocking term — the longest a lower-priority task can hold a resource this task needs. It’s the term everyone omits, and omitting it is why an analysis that passed on paper fails on hardware.

Work the sensor task, which sits below control and Modbus RTU:

C = 200 µs
B = 120 µs   (worst case waiting on the SPI bus manager's mutex)

R₀ = 320
R₁ = 320 + ⌈320/1000⌉·60 + ⌈320/5000⌉·300 = 320 + 60 + 300 = 680
R₂ = 320 + ⌈680/1000⌉·60 + ⌈680/5000⌉·300 = 680     ← converged

680 µs against a 10 ms deadline. Fine, with room. But note that the 120 µs blocking term contributed more than half the initial value — the mutex matters more than the maths people usually do.

Sensor task response time built from 120 microseconds of blocking, 200 of its own work, 60 of control preemption and 300 of Modbus RTU preemption, totalling 680 microseconds.
Response-time analysis, including the blocking term.

Where This Analysis Lies to You

Three ways, and all three bite on an STM32H7.

1. C is not a constant

Rate-monotonic assumes worst-case execution time is a number. On a cached M7 with multiple bus masters, it’s a distribution with an ugly tail.

Your control task measures 60 µs on a quiet bench. Now the Ethernet DMA is writing descriptors into D2 SRAM, the ADC DMA is moving samples, and your task takes a D-cache miss it didn’t take before, followed by a line fill that contends for the AXI bus with the DMA that’s already running. The 60 µs becomes 85 µs, and it does so exactly when the network is busy — which is the load case you didn’t soak-test.

Illustrative execution-time histograms: tight around 60 microseconds on a quiet bench, with a long tail towards 85 microseconds and beyond when DMA traffic causes cache misses.
On a cached M7, execution time has a tail.

Two mitigations, both architectural rather than clever:

Put the hard-deadline path where the variance isn’t. Article 02 placed the control task’s stack, the PID state, and the ADC buffers in DTCM, and the control loop code in ITCM, specifically so cache behaviour never enters the timing analysis for the one task that can’t tolerate it. That’s not an optimisation. It’s what makes C a number instead of a distribution.

And budget headroom for the rest. The 0.72 bound assumes C is honest. If your C values came from a quiet bench, treat 50% as your real ceiling and measure under adversarial load — maximum Ethernet traffic, all sensors polling, an OTA download in flight — before you believe any of it.

2. Sporadic tasks have no period

Modbus RTU frames arrive when the master decides. lwIP wakes when packets show up. Neither has a T.

The fix is to model the minimum inter-arrival time and treat it as a period. For Modbus RTU at 115200 with a 3.5-character inter-frame gap, the master can’t legally send faster than roughly every 5 ms for typical frame sizes — that’s your T.

The danger is when the minimum inter-arrival is a lie. An Ethernet port doesn’t care about your model; a broadcast storm arrives at line rate. Which means the analysis holds only if something enforces the bound, and that something is usually a rate limiter or a bounded queue that drops. Backpressure isn’t a robustness feature. It’s what makes your schedulability analysis true.

3. Self-suspension breaks the model

This is the proof of article 02’s second rule.

Rate-monotonic assumes a task, once running, is only interrupted by higher-priority preemption. A task that voluntarily blocks on I/O in the middle of its execution violates that assumption, and the standard response-time equation stops bounding it correctly.

Practically: put a blocking task high in the priority order and you create a window where a high-priority task is suspended, lower-priority tasks run and take resources, and the high-priority task resumes into contention the analysis never modelled. The response times get worse in a way the arithmetic won’t show you.

Hence the rule. Non-blocking tasks on top, blocking tasks below, and the boundary in the KIC-400 sits between Sensor at 5 and lwIP at 4. Everything above 4 either doesn’t block or blocks only on a bounded, owned resource.


Choosing the Primitive

Semantics first. Performance is a tie-breaker and nothing more.

You need toUseBecause
Move data producer → consumerQueueIt copies. Ownership transfers cleanly.
Say “something happened”, 1:1Task notificationNo queue object, no copy, no extra RAM
Say “something happened”, 1:manyEvent groupMultiple waiters, bit semantics
Protect a resourceMutexOwnership, and priority inheritance
Count available resourcesCounting semaphoreThat’s literally what it counts
Signal from ISR to one taskNotification from ISRFastest path out of interrupt context
Run something laterSoftware timerBut see the gotcha below

Task notifications are faster than queues and use less RAM — no separate object, no copy of the payload. That’s a good reason to prefer them when the semantics already fit: one target task, and you’re signalling an occurrence rather than transferring data. Reaching for a notification because it benchmarked faster, then encoding data into the notification value, produces something that works until you need a second consumer.

Three gotchas worth more than the table

Priority inheritance only works on mutexes. In FreeRTOS, xSemaphoreCreateMutex has it. xSemaphoreCreateBinary does not — a binary semaphore used for mutual exclusion gives you unbounded priority inversion with no warning and no diagnostic. This is the single most common RTOS bug I see in code review, and it’s invisible until a medium-priority task happens to be runnable at the wrong moment.

Timelines comparing a binary semaphore used as a lock, where a medium-priority task causes unbounded inversion, with a mutex where the low-priority holder inherits high priority.
Binary semaphores give no priority inheritance.

Timer callbacks run in the timer service task. All of them, in one context, at configTIMER_TASK_PRIORITY. A callback that blocks — or just takes 5 ms — delays every other timer in the system. Timer callbacks post events. They don’t do work.

All software timer callbacks run sequentially in the timer service task; one 5 ms callback delays the rest.
One slow callback delays every timer.

ISRs above configMAX_SYSCALL_INTERRUPT_PRIORITY cannot call FromISR APIs. They’re not masked by critical sections, which is exactly why you’d put a hard-real-time ISR there, and it means that ISR cannot signal a task. It talks to hardware and to preallocated memory, nothing else. Call xQueueSendFromISR from one and you’ll corrupt kernel state in a way that manifests much later.

Interrupt priorities above configMAX_SYSCALL_INTERRUPT_PRIORITY are never masked and must not call FromISR APIs; those at or below may.
The syscall priority ceiling decides which ISRs may use the kernel.

Mutexes, and Avoiding Them

The best mutex is the one you didn’t need.

Article 02’s config service returns a const snapshot pointer. Readers take no lock, because the data they’re reading is immutable and the writer publishes by swapping the pointer. The control task reads a setpoint with zero blocking time, which means B = 0 for that access, which means it doesn’t appear in the response-time analysis at all.

That pattern — immutable snapshots plus pointer swap, or single-writer with the reader on the same or lower priority — removes more contention than any locking discipline.

When you do need a mutex:

One at a time. A task holding two locks is how deadlock enters a system. If you genuinely need two, define a global lock order, document it in the header, and never take them in the other order.

Hold it for a bounded, short time, and know the number — that number is B in the analysis for every higher-priority task that wants the same resource.

Never block on anything else while holding one. No sockets, no flash writes, no vTaskDelay. The lock’s hold time becomes the other operation’s latency.

Never take one from an ISR. Not “avoid.” Can’t — it’s undefined behaviour, and the ISR has no priority to inherit.

The one giant mutex deserves its own mention because it looks safe: take_global_lock(); do_everything(); give_global_lock();. You’ve serialised the system, converted your RTOS back into a superloop with more overhead, and every task’s B term is now the longest critical section anywhere in the codebase.


Queue Sizing and Backpressure

Queue depth is an architecture decision, and it needs to be made twice — once for the depth, once for what happens when the depth isn’t enough.

Size against the worst-case burst, not the average rate. If the average rate exceeds the drain rate, no depth saves you; the queue only postpones the failure and makes it harder to diagnose. Depth absorbs bursts. It does not fix a throughput mismatch.

For each queue, decide the overflow policy at design time and write it in the header:

QueueDepthOn overflowWhy
ADC samples → Control2Overwrite oldestStale samples are worthless to a control loop
Modbus RTU RX4 framesDrop newestMaster will retry. Protocol handles it.
Fault records32Drop newest, bump a counterNever lose the first fault — it’s the causal one
Log events64Drop oldest, bump a counterRecent context beats old context
Storage requests8Block producer, bounded timeoutConfig writes must not be silently lost

Four different policies across five queues, because the semantics genuinely differ. The fault queue keeping the oldest and the log queue keeping the newest isn’t inconsistency — the first fault is the one that explains the cascade, and the last log line is the one nearest the failure.

Five queues with their depths and overflow policies: overwrite oldest, drop newest, drop newest with a counter, drop oldest with a counter, and block producer with timeout.
Overflow policy is a design decision per queue.

And every drop increments a counter that diagnostics can read. A silent drop is a field bug with no evidence, which is article 03’s anti-pattern #14 arriving through a queue.


Stack Sizing: Measure, Don’t Guess

Stack overflow is the worst class of RTOS bug. It corrupts a neighbouring task’s data and surfaces later, somewhere unrelated, as a fault that makes no sense.

First, the trap that costs people a day: xTaskCreate takes stack depth in words, not bytes. Pass 1024 expecting a 1 KB stack on a Cortex-M and you allocated 4 KB. Nine tasks sized that way is 20-odd KB of RAM you didn’t intend to spend, and the mistake is invisible because everything works.

The method, in order:

Static analysis for the floor. Compile with -fstack-usage, and you get a .su file per translation unit with each function’s frame size. Walk the call graph and sum the deepest path. This catches the error path that never executes on your bench and is therefore invisible to runtime measurement.

Runtime high-water marks for the reality. uxTaskGetStackHighWaterMark() returns the minimum free space ever observed — in words, again. Soak under adversarial load for a week, log it hourly, take the worst.

Take the maximum of both, then add margin. Static analysis misses recursion and function pointers. Runtime measurement misses paths that didn’t run. Neither alone is sufficient, and the pair disagreeing is informative — it usually means an error path with a deep call chain.

Then add the interrupt cost, which isn’t in either number. On a Cortex-M7 with the FPU enabled, an exception frame is 104 bytes when floating-point context is stacked, against 32 bytes without. Nested interrupts stack again. Whether that lands on the task stack or the main stack depends on your configuration, and if you’re using PSP for tasks, it lands on the task’s. Nine tasks × an extra interrupt frame is real RAM.

Stack sizing flow: static analysis and runtime high-water marks, take the maximum plus margin, add the exception frame, and remember xTaskCreate takes words.
Measure the stack two ways, then add the interrupt frame.

Turn on configCHECK_FOR_STACK_OVERFLOW 2 in development, and leave the hook in production writing the offending task name to backup SRAM before it resets. That single record turns an unexplained field reset into a one-line diagnosis.


What to Instrument

Four numbers, logged hourly, that make an RTOS observable instead of mysterious:

Runtime share per task — configGENERATE_RUN_TIME_STATS and uxTaskGetSystemState(). Catches article 03’s God Task while it’s still forming.

Stack high-water mark per task. Trending down means something’s growing.

Peak queue depth, not current. Current depth is almost always zero and tells you nothing; peak tells you whether your burst sizing was right.

Drop counters per queue. A non-zero value that nobody noticed is the most useful thing in a diagnostic dump.

All four fit in well under a hundred bytes and go straight into article 11’s telemetry.


Next

An RTOS gives a task a thread of execution. It says nothing about what that task does over time.

Look at the Modbus RTU task: idle, receiving, processing, responding, back to idle. That’s behaviour, and right now it’s implicit — encoded in where the program counter happens to be sitting inside a sequence of blocking calls. Which works until you need a timeout in the middle of a frame, or a reset during processing, and the state you need to inspect exists only as a stack frame.

Article 06: State Machine Architecture — making behaviour explicit, testable, and independent of who’s executing it.


Series Index

Phase 1 — Architecture
01. Embedded Firmware Architecture Fundamentals
02. Production-Grade Firmware Architecture
03. Firmware Architecture Anti-Patterns

Phase 2 — Implementation
04. HAL vs BSP vs Drivers vs Middleware
05. RTOS Architecture (you are here)
06. State Machine Architecture

Phase 3 — Failure & Resilience
07. Error Handling & Recovery
08. Firmware Security Architecture

Phase 4 — Verification
09. Designing Firmware for Testability

Phase 5 — Product Lifecycle
10. Designing Firmware for 10-Year Products
11. Field Diagnostics & Observability
12. OTA & Safe Firmware Updates

Leave a Reply

Your email address will not be published. Required fields are marked *