Using an RTOS doesn’t make your architecture better. It moves the difficulty.
A system with twenty tasks, fifteen queues, eight mutexes, three event groups and a timer service can be considerably harder to reason about than a superloop, and the failure modes are worse — a superloop that misses a deadline misses it visibly, every cycle. An RTOS that misses one does it under a specific interleaving, once a week, on the unit with the most network traffic.
Article 02 handed you a task table. Nine tasks, nine priorities, nine stack sizes, and two rules stated flatly: priority follows deadline rather than importance, and a task that can block never sits above one that can’t.
I asserted those. This article derives them, using real scheduling analysis on the KIC-400 — and then covers the three places that analysis stops being true on hardware you can buy.
The Concurrency Model Comes First
The question is never “how many tasks.” It’s what independent threads of execution the product actually has, and that’s answerable from requirements before you open an editor.
Five questions define a task. If you can’t answer all five, you don’t have a task, you have a function that hasn’t found its home:
- What wakes it? A timer, a queue, an ISR notification, a socket.
- What’s its deadline? A number, or explicitly none.
- What can block it, and for how long? Bounded, always.
- What does it own? Resources, state, or nothing.
- What happens if it stops? Someone has to notice. Usually the health monitor.
Article 02’s table is those five columns. That’s all a task table is.

Deriving the Task Set
Four reasons justify a separate task. Not “it’s a separate module” — that’s article 03’s task-per-function trap, and it produces forty tasks and 40 KB of stack.
A distinct deadline. Control runs at 1 kHz with a 50 µs jitter budget. Diagnostics runs at 10 Hz with no hard deadline. They cannot share a task, because the shared task inherits the tighter deadline and the looser one’s worst-case execution time.
A distinct blocking behaviour. MQTT blocks inside a TLS handshake for seconds. Modbus RTU must respond within milliseconds. Same task means the Modbus response waits for a handshake.
A distinct failure domain. REQ-06 said a sensor failure must not stop the control loop. Merge sensing into control and that requirement is unsatisfiable — not degraded, unsatisfiable.
A distinct owned resource. The SPI bus manager owns SPI2 and serialises access to it. That serialisation is the task.
Run those four over the KIC-400 requirements and the nine tasks fall out. Two boundary calls worth showing, because they’re the ones that go the other way:
Diagnostics and Health Monitor merged. Both periodic, both low priority, neither blocks, same failure domain. Two tasks would cost 768 bytes of stack and buy nothing. Merged.
Modbus RTU and Modbus TCP kept separate, despite implementing nearly the same register map. Different deadlines (5 ms vs 100 ms), different blocking behaviour (none vs sockets), different failure domains. Same protocol logic, called from two contexts — the code is shared, the execution context isn’t. That distinction is the one people miss when they say “don’t duplicate the Modbus stack.”

Priority Assignment, Properly
Here’s where article 02’s first rule gets its proof.
Rate monotonic, in one paragraph
For periodic tasks with deadlines equal to their periods, the optimal fixed-priority assignment is: shorter period gets higher priority. Not more important. Shorter period. If any fixed-priority assignment can meet all the deadlines, rate-monotonic can — that’s Liu and Layland, 1973, and it’s the whole basis for the rule.
The schedulability bound for n tasks:
U = Σ (Cᵢ / Tᵢ) ≤ n · (2^(1/n) − 1)
For nine tasks that’s 0.72. Under 72% total utilisation, rate-monotonic priorities guarantee every deadline is met. Above it, you need the sharper test below — it might still be fine.
On the KIC-400
Budgets first, measurements after. This is what the table looks like before anyone writes code:
| Task | Period T | Budget C | U = C/T | RM priority |
|---|---|---|---|---|
| Control | 1 ms | 60 µs | 0.060 | 6 (highest) |
| Modbus RTU | 5 ms (sporadic) | 300 µs | 0.060 | 5 |
| Sensor | 10 ms | 200 µs | 0.020 | 5 |
lwIP tcpip_thread | 1 ms (modelled) | 100 µs | 0.100 | 4 |
| Modbus TCP | 50 ms | 800 µs | 0.016 | 3 |
| MQTT | 5 s | 40 ms | 0.008 | 3 |
| Health Monitor | 100 ms | 150 µs | 0.002 | 2 |
| Firmware Update | — | — | — | 2 |
| Logging | — | — | — | 1 |
| Total | 0.266 |
27% against a 72% bound. Comfortable, and the headroom is deliberate — on an M7 with caches you want real margin, for reasons in the next section.

Notice Firmware Update at priority 2 with no period and no utilisation. It’s the most commercially important feature on the list. It has no deadline, so rate-monotonic gives it nothing. That’s article 02’s rule, and now it’s not an opinion.
When the utilisation bound isn’t enough
The bound is sufficient, not necessary — fail it and you might still be schedulable. The exact test is response-time analysis:
Rᵢ = Cᵢ + Bᵢ + Σ ⌈Rᵢ / Tⱼ⌉ · Cⱼ for all j of higher priority
Iterate until it converges. If Rᵢ ≤ Dᵢ, the task meets its deadline.
Bᵢ is the blocking term — the longest a lower-priority task can hold a resource this task needs. It’s the term everyone omits, and omitting it is why an analysis that passed on paper fails on hardware.
Work the sensor task, which sits below control and Modbus RTU:
C = 200 µs
B = 120 µs (worst case waiting on the SPI bus manager's mutex)
R₀ = 320
R₁ = 320 + ⌈320/1000⌉·60 + ⌈320/5000⌉·300 = 320 + 60 + 300 = 680
R₂ = 320 + ⌈680/1000⌉·60 + ⌈680/5000⌉·300 = 680 ← converged
680 µs against a 10 ms deadline. Fine, with room. But note that the 120 µs blocking term contributed more than half the initial value — the mutex matters more than the maths people usually do.

Where This Analysis Lies to You
Three ways, and all three bite on an STM32H7.
1. C is not a constant
Rate-monotonic assumes worst-case execution time is a number. On a cached M7 with multiple bus masters, it’s a distribution with an ugly tail.
Your control task measures 60 µs on a quiet bench. Now the Ethernet DMA is writing descriptors into D2 SRAM, the ADC DMA is moving samples, and your task takes a D-cache miss it didn’t take before, followed by a line fill that contends for the AXI bus with the DMA that’s already running. The 60 µs becomes 85 µs, and it does so exactly when the network is busy — which is the load case you didn’t soak-test.

Two mitigations, both architectural rather than clever:
Put the hard-deadline path where the variance isn’t. Article 02 placed the control task’s stack, the PID state, and the ADC buffers in DTCM, and the control loop code in ITCM, specifically so cache behaviour never enters the timing analysis for the one task that can’t tolerate it. That’s not an optimisation. It’s what makes C a number instead of a distribution.
And budget headroom for the rest. The 0.72 bound assumes C is honest. If your C values came from a quiet bench, treat 50% as your real ceiling and measure under adversarial load — maximum Ethernet traffic, all sensors polling, an OTA download in flight — before you believe any of it.
2. Sporadic tasks have no period
Modbus RTU frames arrive when the master decides. lwIP wakes when packets show up. Neither has a T.
The fix is to model the minimum inter-arrival time and treat it as a period. For Modbus RTU at 115200 with a 3.5-character inter-frame gap, the master can’t legally send faster than roughly every 5 ms for typical frame sizes — that’s your T.
The danger is when the minimum inter-arrival is a lie. An Ethernet port doesn’t care about your model; a broadcast storm arrives at line rate. Which means the analysis holds only if something enforces the bound, and that something is usually a rate limiter or a bounded queue that drops. Backpressure isn’t a robustness feature. It’s what makes your schedulability analysis true.
3. Self-suspension breaks the model
This is the proof of article 02’s second rule.
Rate-monotonic assumes a task, once running, is only interrupted by higher-priority preemption. A task that voluntarily blocks on I/O in the middle of its execution violates that assumption, and the standard response-time equation stops bounding it correctly.
Practically: put a blocking task high in the priority order and you create a window where a high-priority task is suspended, lower-priority tasks run and take resources, and the high-priority task resumes into contention the analysis never modelled. The response times get worse in a way the arithmetic won’t show you.
Hence the rule. Non-blocking tasks on top, blocking tasks below, and the boundary in the KIC-400 sits between Sensor at 5 and lwIP at 4. Everything above 4 either doesn’t block or blocks only on a bounded, owned resource.
Choosing the Primitive
Semantics first. Performance is a tie-breaker and nothing more.
| You need to | Use | Because |
|---|---|---|
| Move data producer → consumer | Queue | It copies. Ownership transfers cleanly. |
| Say “something happened”, 1:1 | Task notification | No queue object, no copy, no extra RAM |
| Say “something happened”, 1:many | Event group | Multiple waiters, bit semantics |
| Protect a resource | Mutex | Ownership, and priority inheritance |
| Count available resources | Counting semaphore | That’s literally what it counts |
| Signal from ISR to one task | Notification from ISR | Fastest path out of interrupt context |
| Run something later | Software timer | But see the gotcha below |
Task notifications are faster than queues and use less RAM — no separate object, no copy of the payload. That’s a good reason to prefer them when the semantics already fit: one target task, and you’re signalling an occurrence rather than transferring data. Reaching for a notification because it benchmarked faster, then encoding data into the notification value, produces something that works until you need a second consumer.
Three gotchas worth more than the table
Priority inheritance only works on mutexes. In FreeRTOS, xSemaphoreCreateMutex has it. xSemaphoreCreateBinary does not — a binary semaphore used for mutual exclusion gives you unbounded priority inversion with no warning and no diagnostic. This is the single most common RTOS bug I see in code review, and it’s invisible until a medium-priority task happens to be runnable at the wrong moment.

Timer callbacks run in the timer service task. All of them, in one context, at configTIMER_TASK_PRIORITY. A callback that blocks — or just takes 5 ms — delays every other timer in the system. Timer callbacks post events. They don’t do work.

ISRs above configMAX_SYSCALL_INTERRUPT_PRIORITY cannot call FromISR APIs. They’re not masked by critical sections, which is exactly why you’d put a hard-real-time ISR there, and it means that ISR cannot signal a task. It talks to hardware and to preallocated memory, nothing else. Call xQueueSendFromISR from one and you’ll corrupt kernel state in a way that manifests much later.

Mutexes, and Avoiding Them
The best mutex is the one you didn’t need.
Article 02’s config service returns a const snapshot pointer. Readers take no lock, because the data they’re reading is immutable and the writer publishes by swapping the pointer. The control task reads a setpoint with zero blocking time, which means B = 0 for that access, which means it doesn’t appear in the response-time analysis at all.
That pattern — immutable snapshots plus pointer swap, or single-writer with the reader on the same or lower priority — removes more contention than any locking discipline.
When you do need a mutex:
One at a time. A task holding two locks is how deadlock enters a system. If you genuinely need two, define a global lock order, document it in the header, and never take them in the other order.
Hold it for a bounded, short time, and know the number — that number is B in the analysis for every higher-priority task that wants the same resource.
Never block on anything else while holding one. No sockets, no flash writes, no vTaskDelay. The lock’s hold time becomes the other operation’s latency.
Never take one from an ISR. Not “avoid.” Can’t — it’s undefined behaviour, and the ISR has no priority to inherit.
The one giant mutex deserves its own mention because it looks safe: take_global_lock(); do_everything(); give_global_lock();. You’ve serialised the system, converted your RTOS back into a superloop with more overhead, and every task’s B term is now the longest critical section anywhere in the codebase.
Queue Sizing and Backpressure
Queue depth is an architecture decision, and it needs to be made twice — once for the depth, once for what happens when the depth isn’t enough.
Size against the worst-case burst, not the average rate. If the average rate exceeds the drain rate, no depth saves you; the queue only postpones the failure and makes it harder to diagnose. Depth absorbs bursts. It does not fix a throughput mismatch.
For each queue, decide the overflow policy at design time and write it in the header:
| Queue | Depth | On overflow | Why |
|---|---|---|---|
| ADC samples → Control | 2 | Overwrite oldest | Stale samples are worthless to a control loop |
| Modbus RTU RX | 4 frames | Drop newest | Master will retry. Protocol handles it. |
| Fault records | 32 | Drop newest, bump a counter | Never lose the first fault — it’s the causal one |
| Log events | 64 | Drop oldest, bump a counter | Recent context beats old context |
| Storage requests | 8 | Block producer, bounded timeout | Config writes must not be silently lost |
Four different policies across five queues, because the semantics genuinely differ. The fault queue keeping the oldest and the log queue keeping the newest isn’t inconsistency — the first fault is the one that explains the cascade, and the last log line is the one nearest the failure.

And every drop increments a counter that diagnostics can read. A silent drop is a field bug with no evidence, which is article 03’s anti-pattern #14 arriving through a queue.
Stack Sizing: Measure, Don’t Guess
Stack overflow is the worst class of RTOS bug. It corrupts a neighbouring task’s data and surfaces later, somewhere unrelated, as a fault that makes no sense.
First, the trap that costs people a day: xTaskCreate takes stack depth in words, not bytes. Pass 1024 expecting a 1 KB stack on a Cortex-M and you allocated 4 KB. Nine tasks sized that way is 20-odd KB of RAM you didn’t intend to spend, and the mistake is invisible because everything works.
The method, in order:
Static analysis for the floor. Compile with -fstack-usage, and you get a .su file per translation unit with each function’s frame size. Walk the call graph and sum the deepest path. This catches the error path that never executes on your bench and is therefore invisible to runtime measurement.
Runtime high-water marks for the reality. uxTaskGetStackHighWaterMark() returns the minimum free space ever observed — in words, again. Soak under adversarial load for a week, log it hourly, take the worst.
Take the maximum of both, then add margin. Static analysis misses recursion and function pointers. Runtime measurement misses paths that didn’t run. Neither alone is sufficient, and the pair disagreeing is informative — it usually means an error path with a deep call chain.
Then add the interrupt cost, which isn’t in either number. On a Cortex-M7 with the FPU enabled, an exception frame is 104 bytes when floating-point context is stacked, against 32 bytes without. Nested interrupts stack again. Whether that lands on the task stack or the main stack depends on your configuration, and if you’re using PSP for tasks, it lands on the task’s. Nine tasks × an extra interrupt frame is real RAM.

Turn on configCHECK_FOR_STACK_OVERFLOW 2 in development, and leave the hook in production writing the offending task name to backup SRAM before it resets. That single record turns an unexplained field reset into a one-line diagnosis.
What to Instrument
Four numbers, logged hourly, that make an RTOS observable instead of mysterious:
Runtime share per task — configGENERATE_RUN_TIME_STATS and uxTaskGetSystemState(). Catches article 03’s God Task while it’s still forming.
Stack high-water mark per task. Trending down means something’s growing.
Peak queue depth, not current. Current depth is almost always zero and tells you nothing; peak tells you whether your burst sizing was right.
Drop counters per queue. A non-zero value that nobody noticed is the most useful thing in a diagnostic dump.
All four fit in well under a hundred bytes and go straight into article 11’s telemetry.
Next
An RTOS gives a task a thread of execution. It says nothing about what that task does over time.
Look at the Modbus RTU task: idle, receiving, processing, responding, back to idle. That’s behaviour, and right now it’s implicit — encoded in where the program counter happens to be sitting inside a sequence of blocking calls. Which works until you need a timeout in the middle of a frame, or a reset during processing, and the state you need to inspect exists only as a stack frame.
Article 06: State Machine Architecture — making behaviour explicit, testable, and independent of who’s executing it.
Series Index
Phase 1 — Architecture
01. Embedded Firmware Architecture Fundamentals
02. Production-Grade Firmware Architecture
03. Firmware Architecture Anti-Patterns
Phase 2 — Implementation
04. HAL vs BSP vs Drivers vs Middleware
05. RTOS Architecture (you are here)
06. State Machine Architecture
Phase 3 — Failure & Resilience
07. Error Handling & Recovery
08. Firmware Security Architecture
Phase 4 — Verification
09. Designing Firmware for Testability
Phase 5 — Product Lifecycle
10. Designing Firmware for 10-Year Products
11. Field Diagnostics & Observability
12. OTA & Safe Firmware Updates