High-performance Java sorting for large datasets.
Project details
A.P.E.X. leverages advanced techniques to efficiently sort large fixed-width 64-bit key/value datasets. With its descriptor-driven approach and support for parallel processing, it optimizes performance while providing an interactive visualizer to understand the sorting execution plan. Ideal for both evaluation and enterprise integration.
A.P.E.X. stands for Adaptive Parallel Extremal Dispatch, a high-performance Java sorting framework tailored for large fixed-width 64-bit key/value record datasets. This innovative framework leverages descriptor-driven radix planning, parallel processing, and advanced techniques like local refinement to efficiently sort unsigned 64-bit keys. A.P.E.X. also supports an optional signed-key ordering mode, ensuring versatility and adaptability in handling different dataset structures.
Adaptive Execution: The A.P.E.X. framework adapts its execution strategy based on the observed input structure. Key parameters such as radix geometry, bucket routes, tuple utilization, and tiny thresholds are optimized dynamically for peak efficiency.
Parallel Processing: Designed for multi-core systems, A.P.E.X. optimizes histogramming, scatter operations, bucket refinement, and workload scheduling to fully utilize available processing power.
Extremal Descriptors: Utilizing per-bucket extremal descriptors, A.P.E.X. employs bitwise operations like OR and AND to minimize unnecessary computations, refining sorting only on the key bits that are unresolved:
VBM = OR ^ AND
Flexible Dispatch: The framework intelligently routes buckets to the most efficient processing path—whether completed, reversed, or refined—thereby optimizing sorting tasks.
A.P.E.X. processes packed records, where each record consists of an 8-byte key and an 8-byte value, resulting in 16 bytes per record. The key is sorted while maintaining its associated value, often representing a row ID or an index pointer.
The A.P.E.X. repository includes:
To utilize A.P.E.X., ensure the following requirements are met:
A.P.E.X. can be invoked as a library as shown below:
import java.lang.foreign.Arena;
import io.github.strmckr.apex.Apex;
import io.github.strmckr.apex.Configuration;
import io.github.strmckr.apex.generator.DataMode;
try (Arena arena = Arena.ofShared()) {
Apex.Data data = Apex.Data.generated(DataMode.RANDOM, 1_000_000L, arena);
Apex.Result result = Apex.sortAndVerify(data, Configuration.defaultConfiguration());
}
The CLI can also be employed for sorting tasks:
java --enable-preview --enable-native-access=ALL-UNNAMED --add-modules jdk.incubator.vector -Xmx16G -XX:MaxDirectMemorySize=80g -jar A.P.E.X/target/apex-1.0.0-SNAPSHOT.jar mode=RANDOM records=10m threads=16
For further information, visit the live visualizer and access the complete documentation for detailed instructions and guides on maximizing A.P.E.X.'s potential.
Comments
0Start the conversation
Share the first comment.