Asynchronous Processing: Queues, Pub/Subs and Streams
Let's start a review on arguably the most simple-looking thing that most students overlook, but deceptively complicated at scale, and compare the most common methods for going about it.
Synchronous: You run A, wait for it to complete, then run B, wait for it to complete, then C. This is also called blocking.
Asynchronous: You run A, B, and C without waiting for each one to finish before starting the next. All run independently. This is also called non-blocking.
When should Asynchronous Processing be used?
Assume ONE core for simplicity:
BAD fit (CPU tasks):
Example: A (Compress a video), B (Calculate frames), C (Encrypt a file).
Reason: Every task actively consumes CPU cycles to calculate and do things, etc.. Trying to run them asynchronously on one core adds extra time needed for context switching without speeding anything up. Sequential is faster in this case.
GOOD fit (I/O tasks):
Example: A (Wait for user upload), B (Transform data in memory), C (Send a network request).
Reason: A and C spend most of their time “idling”, waiting on the network or hardware inputs. Instead of letting the CPU sit doing nothing, asynchronous processing gives control to task B so it can run while A and C wait. (Remember Time-slicing in Process Management Mr. Nam has shared)
This extends out to many cores, many machines and many microservices.
Asynchronous vs Parallelism
Given three dishes: A, B, C.
Synchronous (1 person, and blocking processing):
You put Dish A in the oven and stand in front of the oven until it finishes. Then you start Dish B.
It is bad, because why are you standing in front of the oven when there are B and C to make?
Asynchronous (1 person, non-blocking processing):
You put Dish A in the oven. While it bakes (waiting on I/O), you move over to prep Dish B. When the oven beeps, you retrieve Dish A.
It is efficient for the case of one person. You don’t do nothing when something needs to be waited.
Parallelism (3 people):
Person 1 cooks A, Person 2 cooks B, and Person 3 cooks C at the same time.
Increases efficiency via hardware (or more people, or more cores, or more services, etc.)
Asynchronous Processing across Services
Message Queues
What:
A queue that holds messages.
Each message is processed by one consumer. After its picked up and processed (via an Acknowledgement), the message is removed from the queue.
Can be understood as Single Producer - Single Consumer model.
When:
Good fit for sending tasks to another service. Could be used to ask the Email Service to send an email, or to ask the Logging Service to log a certain event, etc.
Example: AWS Simple Queue Service ($25 each month for AWS is still too low for Mr Nam, he still got more to spare), RabbitMQ.
Easy to misunderstand stuff:
A message can be sent to many queues. It does not require exact 1 to 1, just point to point. In RabbitMQ, it’s also known as “routing”.
Yes, using a Goroutine (Go), or a Future (Java), or even a Promise (JS) works still, only if the required services sit on the same stack. Messaging allows asynchronous across stacks.
Publisher and Subscriber
What:
A model where subscribers are notified when a source (publisher) announces into a subscribed topic.
Can be understood as Many Producers - Many Consumers model.
A model where messages are kept in a stream, and persisted.
When:
Good fit high-throughput data streams, audit logs and real-time stuff.
Example: Kafka.
In a nutshell
Message Queues
Pub/Sub
Event Streams
Model
One to one
One to many
One to many
Node Knowledge
Publishers know which consumers will take it. Consumers don’t know which publishers will send it.
Publishers and consumers don’t know each other.
Full
Persistence
Until consumed or acknowledged.
Transient (If a subscriber wasn’t listening at the time of broadcast, it would miss it)
Long term.
Message/Packet Tracking
The Message Broker (RabbitMQ, etc. itself) tracks whether consumers acknowledged the message.
Transient (Fire and forget type)
Messages are logged. Consumers track their progress.
Primary Use Case
Task Offloading
Event-based Actions
Historical logging
Simple analogy (Nam, Minh, Phong, MTuan, HTuan):
Minh logs a ticket to fix into Frontend’s queue. MTuan and HTuan looks at Frontend queue. MTuan grabs the ticket, removes it from the queue (The message is consumed). HTuan is still waiting, no tickets to do. This is a Message Queue. MTuan and HTuan are consumers. Minh is a producer.
Nam and HTuan subscribe to “Codex” and “How to give Anh money”. Phong and MTuan subscribe to “Claude” and “Codex”. Minh broadcasts a message to “Codex”, all 4 hear the message. Minh broadcasts a message to “Claude”, only Phong and MTuan hear the message. This is a Pub/Sub.
Nam, Phong, MTuan and HTuan are having a meeting. Minh is also there, but not speaking anything, just screen-recording and noting. Minh is making an event stream. Later on, when Nam wants to see how much money Phong volunteered to give Anh, Nam can read the notes and replay the meeting.
Asynchronous in Monoliths
Asynchronous is about being as efficient as possible by minimizing idle time waiting for things that could be used for actual work. For further and clearer definitions of cloud-based resources, check out Appendix A.
Assume each PSQL read takes 5s. Assume each computation takes 2s. Request A wants to read data. Request B wants to compute. Request C wants to read data.
Synchronous:
[0s] Request A, B and C come in sequentially.
[0s] The server starts with A. The server calls PSQL to read and waits.
[5s] PSQL returns. The server responds to A. A finishes (total 5s). The server continues with B.
[7s] The server finishes computing and responds to B. B finishes (total 7s). The server continues with C. The server calls PSQL to read and waits.
[12s] The server responds to C. C finishes (total 12s).
Asynchronous:
[0s] Request A, B and C come in sequentially.
[0s] The server starts with A. It calls PSQL to read, but puts A on the background. It starts with B.
[2s] The server finishes computing and responds to B. B finishes (total 2s). The server starts with C. It calls PSQL to read, and puts C on the background. It idles.
[5s] PSQL returns for A. The server responds to A. A finishes (total 5s). PSQL starts for C. The server idles.
[10s] PSQL returns for C. The server responds to C. C finishes (total 10s).
So, by just simply being asynchronous, we have saved 2 seconds until completion, and total response time (before: 24s, after: 17s) shaved by 7s whole.
The Problems of Going Async
Appendix
Appendix A. CPUs, Cores, and Cloud-based vCPUs and vCores
It seems very easy to mistake these things, when it’s advertised on cloud plans such as AWS, GCP or other providers. AEStore stack currently uses 2 shared vCPUs.
A CPU: The physical chip that is plugged into the motherboard. A CPU can have many Cores.
A Core: The physical computing component of a CPU, includes its own registers set, ALU, and such.
A Thread: Usually in this context, referred to the Hardware Thread. A Core has 2 hardware threads. This technique is called Hyperthreading, or SMT, which the core duplicates its architecture logically, so that the OS thinks that there are 2 cores, while there is only 1 physical core.
A vCPU or a vCore (seems to be used the same): The hardware thread. It does not inherently mean when you buy 2 vCPUs, you get 2 threads, as there is a “subscription ratio”. 1 thread can be split into 4 vCPUs sold to 4 customers, if you need to run something, but one person is already running it, this will get tracked as “CPU Steal%”. A “dedicated” vCPU guarantees 1 mapping in most cases.