Configuration
Arneb loads configuration from three sources with the following precedence (highest wins):
- CLI arguments (
--port,--config,--role) - Environment variables (
ARNEB_PORT,ARNEB_BIND_ADDRESS, etc.) - Configuration file (
arneb.toml) - Built-in defaults
Configuration File
By default, Arneb looks for arneb.toml in the current directory. Specify a different path with:
cargo run --bin arneb -- --config /path/to/config.tomlServer Settings
| Field | Type | Default | Env Var | Description |
|---|---|---|---|---|
bind_address | string | "127.0.0.1" | ARNEB_BIND_ADDRESS | Address to bind the server to |
port | integer | 5432 | ARNEB_PORT | PostgreSQL wire protocol port |
max_worker_threads | integer | (CPU count) | ARNEB_MAX_WORKER_THREADS | Maximum worker threads for query execution |
max_memory_mb | integer | (system dependent) | ARNEB_MAX_MEMORY_MB | Maximum memory in MB |
Ports
| Service | Port | Roles |
|---|---|---|
| pgwire (PostgreSQL protocol) | port | standalone, coordinator |
| Trino client protocol (HTTP) | [trino] port (default 8080) | standalone, coordinator |
| Web UI | port + 1000 | standalone, coordinator |
| Flight RPC | 9090 | all roles |
Trino Client Protocol
| Field | Type | Default | Env Var | CLI | Description |
|---|---|---|---|---|---|
[trino] enabled | bool | true | ARNEB_TRINO_ENABLED | --no-trino | Serve the Trino client REST protocol |
[trino] port | integer | 8080 | ARNEB_TRINO_PORT | --trino-port | HTTP port of the Trino listener |
See Trino Client Compatibility for what the listener supports.
Tuning Knobs: Build-Time vs Runtime
Arneb separates configuration into two classes. Knowing which is which keeps builds reproducible and avoids silent misconfiguration.
Runtime-tunable — anything that can change without recompiling (per-node memory budget, parallelism, log level, allocator decay). These are exposed as ARNEB_* environment variables / arneb.toml fields, follow the precedence above, and the effective value is logged at startup. Override freely per deployment or experiment.
Build-time — anything that can only be decided at compile/link time (Cargo features, codegen flags). These live in version-controlled build config (Cargo.toml, .cargo/config.toml) and are changed in source, never via an environment variable. A binary's behaviour must be a pure function of its committed inputs.
Two rules follow:
- Never override a build-time parameter with an environment variable. To experiment with a build-time setting, use an explicit
cargo --features/--profileinvocation, then commit the chosen default. - Never rely on a third-party allocator/runtime's own environment variable (e.g. jemalloc's
MALLOC_CONF/_RJEM_MALLOC_CONF). Those are silent on typos and prefix mismatches. Every runtime knob goes through anARNEB_*variable that Arneb reads, applies, and logs — so a wrong value is visible, not silently ignored.
Memory / Allocator Tuning
Arneb uses jemalloc and returns freed pages to the OS promptly so the cgroup memory peak reflects the engine's true working set, not allocator history. The page-decay interval is runtime-tunable and set in-binary at startup (default shown), with the effective value logged.
| Knob | Default | Env Var | Description |
|---|---|---|---|
| dirty/muzzy page decay | 500 ms | ARNEB_DIRTY_DECAY_MS | How long jemalloc holds freed pages before madvise-ing them back to the OS. Lower = tighter RSS but more page re-faults; 0 returns immediately (slowest); higher lets RSS drift up. 500 ms is the measured sweet spot. Applied via mallctl — do not set MALLOC_CONF, it is ignored. |
| spill budget | config / cgroup | ARNEB_SPILL_BUDGET_BYTES | Per-node budget (bytes) a spillable operator (SemiJoin/HashJoin build) reserves against before spilling to disk. Overrides the [memory] spill_budget_bytes config field. Exists as an env var because the bench config is COPYed into the docker image at build time — the env override retunes without an image rebuild. |
| query memory cap | config | ARNEB_QUERY_MAX_MEMORY_BYTES | Per-task cumulative allocation cap (bytes). When the query's tracked MemoryReservation crosses it, the query fails cleanly with ResourceExhausted instead of OOM-killing the worker. Overrides [memory] query_max_memory_per_node. |
The startup log line memory pool installed … spill_budget_bytes=… spill_budget_source=… confirms the resolved value and its source (env / config / cgroup_v2 / cgroup_v1 / unbounded); a source=env line confirms the override took effect.
Table Registration
Register tables directly in the config file:
[[tables]]
name = "lineitem"
path = "/data/lineitem.parquet"
format = "parquet"
[[tables]]
name = "orders"
path = "/data/orders.csv"
format = "csv"| Field | Type | Required | Description |
|---|---|---|---|
name | string | yes | Table name used in SQL queries |
path | string | yes | File path (local or remote s3://, gs://, az://) |
format | string | yes | File format: "parquet" or "csv" |
Object Store Configuration
S3
[storage.s3]
region = "us-east-1"
endpoint = "http://localhost:9000" # For MinIO/LocalStack; omit for AWS
allow_http = true # Required when endpoint uses HTTP
# access_key_id = "minioadmin" # Optional: falls back to env/IAM
# secret_access_key = "minioadmin"Credential precedence: config file → AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY env vars → IAM role / instance profile.
GCS
[storage.gcs]
service_account_path = "/path/to/service-account.json"Catalog Configuration
Register external catalogs (e.g., Hive Metastore; its Iceberg tables are read through the same catalog):
[[catalogs]]
name = "datalake"
type = "hive"
metastore_uri = "127.0.0.1:9083"
default_schema = "default"
# Per-catalog storage override (merges with global [storage])
[catalogs.storage.s3]
region = "us-east-1"
endpoint = "http://localhost:9000"
allow_http = true| Field | Type | Required | Description |
|---|---|---|---|
name | string | yes | Catalog name (used as first part of catalog.schema.table) |
type | string | yes | Catalog type (currently "hive"; Iceberg tables in the metastore are included, see Iceberg) |
metastore_uri | string | yes | host:port of the Hive Metastore (no scheme prefix) |
default_schema | string | no | Default schema within the catalog |
Cluster Configuration
For distributed mode (worker nodes):
[cluster]
rpc_port = 9091
coordinator_address = "127.0.0.1:9090"
worker_id = "worker-1"| Field | Type | Required | Description |
|---|---|---|---|
rpc_port | integer | yes | Flight RPC port for this worker |
coordinator_address | string | yes | host:port of the coordinator's Flight RPC |
worker_id | string | yes | Unique identifier for this worker |
See Distributed Mode for full setup instructions.
Authentication
[auth] is the server's client authentication: one credential store intended for every client-facing listener. Today it is enforced on the pgwire port; the Trino listener will verify HTTP Basic passwords against the same [[auth.users]] in a follow-up (see Limitations).
By default every connection is accepted without a password (type = "none"), so existing setups keep working. To require passwords, add an [auth] section:
[auth]
type = "password" # "none" (default) | "password"
[[auth.users]]
name = "alice"
password_hash = "SCRAM-SHA-256$4096:Mzusu25I7rTxxYUaRIIPag==$qQmyyOdRet5ULfbhSwiX8OP4QInk+Monhm0zSkWcgGs=:SgOFTmuXsLnO/AHKLu3a+BAiJ0Yx3JcIj/wNsJX1MxY="| Field | Type | Required | Description |
|---|---|---|---|
type | string | no | "none" (default) or "password" |
users[].name | string | yes | Login name, matched against the client's user |
users[].password_hash | string | yes | SCRAM-SHA-256 verifier (see below) |
password mode uses SCRAM-SHA-256, the same mechanism as PostgreSQL's password_encryption = scram-sha-256. psql, JDBC, psycopg2/3, DBeaver, tokio-postgres and other modern clients support it. The password never crosses the network, and the config file holds only a salted verifier in PostgreSQL's pg_authid.rolpassword format, not the password. A verifier copied from PostgreSQL (SELECT rolpassword FROM pg_authid WHERE rolname = '…') works as-is. To generate one:
echo -n 'my-password' | arneb hash-password # or run it and type at the prompt
echo -n 'my-password' | arneb hash-password --iterations 100000--iterations sets the PBKDF2 iteration count stored in the verifier. The default, 4096, matches PostgreSQL. For production, use a much higher count (for example 100000 or more) to slow down offline guessing if the config file leaks. The cost is paid by the client on every connection, so measure connection latency for your clients before going very high.
Behavior:
- A wrong password or an unknown user is rejected with
FATAL 28P01 password authentication failed for user "…". Unknown users get the same challenge as real ones, so an attacker can't tell which user names exist. - One startup handler serves both the Simple and Extended Query protocols, so every client path gets the same check.
- The effective mode is logged at startup on the
arneb::configtarget (pgwire authentication effective mode auth="password (scram-sha-256)" users=N). - Invalid
[auth]config stops startup. This covers an unknowntype, a malformed hash, duplicate or empty user names,type = "password"with no users, and a plaintextpasswordkey.
Limitations: [auth] is enforced only on the pgwire port. These listeners still accept anyone, so keep them on a trusted network:
- the Trino client protocol port (
[trino] port, default8080, enabled by default): anyX-Trino-Useris accepted without a password. Startup logs a WARN onarneb::configwhentype = "password"is set and this listener is on. Where password protection is required, disable it with--no-trinoorARNEB_TRINO_ENABLED=falseuntil the follow-up (HTTP Basic over TLS, checked against the same SCRAM verifiers) lands; - the Web UI (pgwire port + 1000);
- the Flight RPC port between the coordinator and its workers.
The pgwire port does not support TLS yet. SCRAM keeps the password itself off the wire, but query traffic is still plaintext. Channel binding (SCRAM-SHA-256-PLUS) is not offered.
CLI Arguments
arneb [OPTIONS] [COMMAND]
Commands:
hash-password Read a password from stdin and print a SCRAM-SHA-256
verifier for [[auth.users]] password_hash
Options:
--config <PATH> Path to configuration file
--port <PORT> Override the pgwire port
--role <ROLE> Server role: standalone, coordinator, or workerExample: Standalone with Local Files
bind_address = "127.0.0.1"
port = 5432
[[tables]]
name = "lineitem"
path = "/data/tpch/lineitem.parquet"
format = "parquet"
[[tables]]
name = "orders"
path = "/data/tpch/orders.parquet"
format = "parquet"Example: Distributed with Hive Catalog
bind_address = "0.0.0.0"
port = 5432
[storage.s3]
region = "us-east-1"
endpoint = "http://minio:9000"
allow_http = true
[[catalogs]]
name = "datalake"
type = "hive"
metastore_uri = "hive-metastore:9083"
default_schema = "default"
[catalogs.storage.s3]
region = "us-east-1"
endpoint = "http://minio:9000"
allow_http = true