Fork/Join framework and parallel streams
The Fork/Join framework (Java 7) runs divide-and-conquer work on all CPU cores:
- Write a RecursiveTask<V> (returns a value) or RecursiveAction (no value) with a compute() method.
- If the task is small enough (below a threshold), compute it directly. Otherwise split it, fork() one half, compute the other, and join() the results.
- A ForkJoinPool runs the tasks. Each worker has its own queue, and idle workers steal tasks from busy ones, so all cores stay busy.
Parallel streams use the same machinery: list.parallelStream() or stream.parallel() splits the source and processes the pieces in the shared ForkJoinPool.commonPool() (one worker fewer than the number of cores).
Parallel isn't automatically faster. It helps for large, CPU-bound work on sources that split cheaply (arrays, ArrayList, IntStream.range) with independent, stateless operations. It hurts for small inputs, LinkedList or I/O sources, blocking calls and anything that shares mutable state.
Example
class SumTask extends RecursiveTask<Long> {
private static final int THRESHOLD = 10_000;
private final long[] numbers;
private final int from, to;
SumTask(long[] numbers, int from, int to) {
this.numbers = numbers; this.from = from; this.to = to;
}
@Override
protected Long compute() {
if (to - from <= THRESHOLD) { // small enough: just loop
long sum = 0;
for (int i = from; i < to; i++) sum += numbers[i];
return sum;
}
int mid = (from + to) >>> 1;
SumTask left = new SumTask(numbers, from, mid);
SumTask right = new SumTask(numbers, mid, to);
left.fork(); // run the left half asynchronously
long rightSum = right.compute(); // do the right half in this thread
return left.join() + rightSum; // wait for (or help with) the left half
}
}
long[] data = LongStream.rangeClosed(1, 10_000_000).toArray();
long total = ForkJoinPool.commonPool().invoke(new SumTask(data, 0, data.length)); // 50000005000000
long same = LongStream.rangeClosed(1, 10_000_000).parallel().sum(); // same resultList<Integer> results = new ArrayList<>();
IntStream.range(0, 1_000).parallel().forEach(results::add); // WRONG: ArrayList isn't thread-safe
List<Integer> safe = IntStream.range(0, 1_000).parallel().boxed().toList(); // let the stream collect
orders.parallelStream().map(o -> callPaymentApi(o)).toList(); // blocking calls clog the shared common pool
names.parallelStream().findFirst(); // order-preserving: slower than findAny() in parallelCommon mistake
Adding .parallel() to make code faster without measuring. For small collections or blocking work it's usually slower, and with shared mutable state it's wrong.
Under the hood
Measure before parallelising, ideally with JMH, because splitting, scheduling and merging have real costs. Order matters: findFirst, forEachOrdered and limit force ordering work that findAny, forEach and unordered() avoid. Since every parallel stream in the JVM shares the common pool, one slow or blocking task can starve all the others; for blocking I/O use virtual threads or a dedicated executor instead.
Check yourself
Which pool runs parallel streams by default?
How this connects
Know these first
Where this leads
You've reached the end of this thread. Try a learning path for what's next.
Part of Multithreading: beginner to advanced.
Was this lesson helpful?
Finished reading? Mark it complete to track your progress.