Module 8

Introduction to Pthreads

With MPI, we launched several processes, each with its own memory. Now we will use threads: several paths of execution inside one process. They can read the same variables and arrays. That makes some kinds of cooperation easier, but it also means we must be careful when two threads change the same data.

This module follows the opening of Chapter 4 of Pacheco and Malensek, especially the hello program in §4.2 and the divided-work pattern in §4.3. We will build two smaller programs together:

Processes and threads

An MPI process has its own ordinary variables. To send a value to another process, we used an MPI call. Threads in one process can access the same global variables and arrays directly:

one process on one PicoCluster board
  main thread ─┬─ worker 0
               ├─ worker 1       shared arrays and global variables
               ├─ worker 2
               └─ worker 3

Each thread also has its own function calls and local variables. A Pthreads program runs on one machine; creating four threads does not place one thread on each of four PicoCluster boards. The main thread exists before we create any workers. With four workers, there are five threads in this program until the workers finish.

Predict: If worker 0 changes a shared array element, can worker 1 read that change? Would the same statement be true of an ordinary array in two separate MPI processes?

First program: say hello

Create pth_hello.c. The complete file is available here.

#include <stdio.h>
#include <pthread.h>

#define THREAD_COUNT 4

void *hello(void *argument) {
    long id = *(long *)argument;
    printf("Hello from worker %ld\n", id);
    return NULL;
}

int main(void) {
    pthread_t workers[THREAD_COUNT];
    long ids[THREAD_COUNT];

    printf("Main thread: starting %d workers\n", THREAD_COUNT);
    for (long t = 0; t < THREAD_COUNT; t++) {
        ids[t] = t;
        pthread_create(&workers[t], NULL, hello, &ids[t]);
    }

    for (int t = 0; t < THREAD_COUNT; t++) {
        pthread_join(workers[t], NULL);
    }
    printf("Main thread: all workers finished\n");
    return 0;
}

Compile on PicoCluster with gcc and the -pthread option, then run the executable directly:

gcc -Wall -pthread -o pth_hello pth_hello.c
./pth_hello

-pthread tells the compiler and linker to use the thread support needed by this program. On our current PicoCluster installation, these examples also compile without -pthread. We keep it in the commands as the usual portable compiler option for Pthreads programs. There is no mpirun here: ./pth_hello starts one process, and that process creates its workers.

The four worker lines may appear in any order. The final line appears only after all four pthread_join calls return. Run it several times: do the workers always greet in the same order?

Piece What it does
pthread_t workers[THREAD_COUNT] Holds the handles we need to join the workers later.
pthread_create(&workers[t], NULL, hello, &ids[t]) Starts a worker in the hello function and passes it the address of its own ID.
void *hello(void *argument) A thread function takes one generic pointer argument and returns a generic pointer.
long id = *(long *)argument Treats that argument as a pointer to long and reads the ID stored there.
pthread_join(workers[t], NULL) Waits for that worker to finish.

Read the argument-passing lines as a pair: &ids[t] sends the address of this worker’s ID; *(long *)argument reads the number at that address. The Pthreads interface uses void * so the same interface can accept the address of many kinds of data. For these examples, the data is just one long.

The ids array matters. Each worker receives a different, stable array element. Passing &t, the address of the changing loop variable, would not reliably give each worker its intended ID. These IDs are values we assign; Pthreads does not give us MPI-style ranks automatically.

Discuss: What might happen if we removed the join loop and allowed main to return immediately? Why can the final line be trusted to come after all worker greetings, even though the greetings have no fixed order?

Second program: divide a sum

We have already divided a sum across MPI processes. This time, four threads will add different consecutive blocks of the integers from 1 through 10,000,000. A worker keeps its running total in a local variable, then writes one distinct slot of a shared partial array. After joining, the main thread combines those four slots.

Worker ID Integers to add Shared result slot
0 1 through 2,500,000 partial[0]
1 2,500,001 through 5,000,000 partial[1]
2 5,000,001 through 7,500,000 partial[2]
3 7,500,001 through 10,000,000 partial[3]

Predict: Does any integer appear in two blocks? Does any worker need to read another worker’s partial result?

Start pth_sum.c with these shared declarations and the worker function:

#define THREAD_COUNT 4
#define N 10000000LL

long long partial[THREAD_COUNT];

void *add_block(void *argument) {
    long id = *(long *)argument;
    long long block_size = N / THREAD_COUNT;
    long long first = id * block_size + 1;
    long long last = (id + 1) * block_size;
    long long local_sum = 0;

    for (long long i = first; i <= last; i++) {
        local_sum += i;
    }
    partial[id] = local_sum;
    return NULL;
}

first, last, and local_sum are local to the worker’s function call. The partial array is shared. Each worker writes a different element, so this version does not require a lock. We chose an N divisible by THREAD_COUNT, keeping the block arithmetic simple.

In main, create the workers just as in the hello program. Join all of them before reading partial:

long long total = 0;
for (int t = 0; t < THREAD_COUNT; t++) {
    pthread_join(workers[t], NULL);
    total += partial[t];
}
printf("Sum from 1 through %lld: %lld\n", N, total);

The complete program includes the creation loop and prints each partial result in ID order after the joins. Compile and run:

gcc -Wall -pthread -o pth_sum pth_sum.c
./pth_sum

The final sum should be 50,000,005,000,000. You can check it using \(N(N+1)/2\) with \(N=10{,}000{,}000\). The program uses long long for the totals because the answer is much larger than a 32-bit int can hold.

Try: Change THREAD_COUNT to 2 or 5, recompile, and run again. The answer should stay the same. More threads do not automatically mean a faster run: creating and coordinating threads has a cost, and this example is here to make ownership of the work clear.

Why not use one shared total?

It is tempting to replace each worker’s local_sum and partial[id] with total += i in the loop. That would make several threads read and write the same variable without coordination. The result would not be reliable. Chapter 4 uses a similar shared-sum problem to motivate synchronization. In later modules, we will investigate the bug and learn when to use mutexes and other tools.

For now, remember the safe pattern we used: separate work, separate partial results, join, then combine. It resembles our MPI reductions, but the partial array is already in the same process’s memory; we do not send messages between these workers.

Back to top