Skip to content

Configuration ​

Arneb loads configuration from three sources with the following precedence (highest wins):

  1. CLI arguments (--port, --config, --role)
  2. Environment variables (ARNEB_PORT, ARNEB_BIND_ADDRESS, etc.)
  3. Configuration file (arneb.toml)
  4. Built-in defaults

Configuration File ​

By default, Arneb looks for arneb.toml in the current directory. Specify a different path with:

bash
cargo run --bin arneb -- --config /path/to/config.toml

Server Settings ​

FieldTypeDefaultEnv VarDescription
bind_addressstring"127.0.0.1"ARNEB_BIND_ADDRESSAddress to bind the server to
portinteger5432ARNEB_PORTPostgreSQL wire protocol port
max_worker_threadsinteger(CPU count)ARNEB_MAX_WORKER_THREADSMaximum worker threads for query execution
max_memory_mbinteger(system dependent)ARNEB_MAX_MEMORY_MBMaximum memory in MB

Ports ​

ServicePortRoles
pgwire (PostgreSQL protocol)portstandalone, coordinator
Trino client protocol (HTTP)[trino] port (default 8080)standalone, coordinator
Web UIport + 1000standalone, coordinator
Flight RPC9090all roles

Trino Client Protocol ​

FieldTypeDefaultEnv VarCLIDescription
[trino] enabledbooltrueARNEB_TRINO_ENABLED--no-trinoServe the Trino client REST protocol
[trino] portinteger8080ARNEB_TRINO_PORT--trino-portHTTP port of the Trino listener

See Trino Client Compatibility for what the listener supports.

Tuning Knobs: Build-Time vs Runtime ​

Arneb separates configuration into two classes. Knowing which is which keeps builds reproducible and avoids silent misconfiguration.

Runtime-tunable — anything that can change without recompiling (per-node memory budget, parallelism, log level, allocator decay). These are exposed as ARNEB_* environment variables / arneb.toml fields, follow the precedence above, and the effective value is logged at startup. Override freely per deployment or experiment.

Build-time — anything that can only be decided at compile/link time (Cargo features, codegen flags). These live in version-controlled build config (Cargo.toml, .cargo/config.toml) and are changed in source, never via an environment variable. A binary's behaviour must be a pure function of its committed inputs.

Two rules follow:

  • Never override a build-time parameter with an environment variable. To experiment with a build-time setting, use an explicit cargo --features / --profile invocation, then commit the chosen default.
  • Never rely on a third-party allocator/runtime's own environment variable (e.g. jemalloc's MALLOC_CONF / _RJEM_MALLOC_CONF). Those are silent on typos and prefix mismatches. Every runtime knob goes through an ARNEB_* variable that Arneb reads, applies, and logs — so a wrong value is visible, not silently ignored.

Memory / Allocator Tuning ​

Arneb uses jemalloc and returns freed pages to the OS promptly so the cgroup memory peak reflects the engine's true working set, not allocator history. The page-decay interval is runtime-tunable and set in-binary at startup (default shown), with the effective value logged.

KnobDefaultEnv VarDescription
dirty/muzzy page decay500 msARNEB_DIRTY_DECAY_MSHow long jemalloc holds freed pages before madvise-ing them back to the OS. Lower = tighter RSS but more page re-faults; 0 returns immediately (slowest); higher lets RSS drift up. 500 ms is the measured sweet spot. Applied via mallctl — do not set MALLOC_CONF, it is ignored.
spill budgetconfig / cgroupARNEB_SPILL_BUDGET_BYTESPer-node budget (bytes) a spillable operator (SemiJoin/HashJoin build) reserves against before spilling to disk. Overrides the [memory] spill_budget_bytes config field. Exists as an env var because the bench config is COPYed into the docker image at build time — the env override retunes without an image rebuild.
query memory capconfigARNEB_QUERY_MAX_MEMORY_BYTESPer-task cumulative allocation cap (bytes). When the query's tracked MemoryReservation crosses it, the query fails cleanly with ResourceExhausted instead of OOM-killing the worker. Overrides [memory] query_max_memory_per_node.

The startup log line memory pool installed … spill_budget_bytes=… spill_budget_source=… confirms the resolved value and its source (env / config / cgroup_v2 / cgroup_v1 / unbounded); a source=env line confirms the override took effect.

Table Registration ​

Register tables directly in the config file:

toml
[[tables]]
name = "lineitem"
path = "/data/lineitem.parquet"
format = "parquet"

[[tables]]
name = "orders"
path = "/data/orders.csv"
format = "csv"
FieldTypeRequiredDescription
namestringyesTable name used in SQL queries
pathstringyesFile path (local or remote s3://, gs://, az://)
formatstringyesFile format: "parquet" or "csv"

Object Store Configuration ​

S3 ​

toml
[storage.s3]
region = "us-east-1"
endpoint = "http://localhost:9000"   # For MinIO/LocalStack; omit for AWS
allow_http = true                     # Required when endpoint uses HTTP
# access_key_id = "minioadmin"       # Optional: falls back to env/IAM
# secret_access_key = "minioadmin"

Credential precedence: config file → AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY env vars → IAM role / instance profile.

GCS ​

toml
[storage.gcs]
service_account_path = "/path/to/service-account.json"

Catalog Configuration ​

Register external catalogs (e.g., Hive Metastore; its Iceberg tables are read through the same catalog):

toml
[[catalogs]]
name = "datalake"
type = "hive"
metastore_uri = "127.0.0.1:9083"
default_schema = "default"

# Per-catalog storage override (merges with global [storage])
[catalogs.storage.s3]
region = "us-east-1"
endpoint = "http://localhost:9000"
allow_http = true
FieldTypeRequiredDescription
namestringyesCatalog name (used as first part of catalog.schema.table)
typestringyesCatalog type (currently "hive"; Iceberg tables in the metastore are included, see Iceberg)
metastore_uristringyeshost:port of the Hive Metastore (no scheme prefix)
default_schemastringnoDefault schema within the catalog

Cluster Configuration ​

For distributed mode (worker nodes):

toml
[cluster]
rpc_port = 9091
coordinator_address = "127.0.0.1:9090"
worker_id = "worker-1"
FieldTypeRequiredDescription
rpc_portintegeryesFlight RPC port for this worker
coordinator_addressstringyeshost:port of the coordinator's Flight RPC
worker_idstringyesUnique identifier for this worker

See Distributed Mode for full setup instructions.

Authentication ​

[auth] is the server's client authentication: one credential store intended for every client-facing listener. Today it is enforced on the pgwire port; the Trino listener will verify HTTP Basic passwords against the same [[auth.users]] in a follow-up (see Limitations).

By default every connection is accepted without a password (type = "none"), so existing setups keep working. To require passwords, add an [auth] section:

toml
[auth]
type = "password"   # "none" (default) | "password"

[[auth.users]]
name = "alice"
password_hash = "SCRAM-SHA-256$4096:Mzusu25I7rTxxYUaRIIPag==$qQmyyOdRet5ULfbhSwiX8OP4QInk+Monhm0zSkWcgGs=:SgOFTmuXsLnO/AHKLu3a+BAiJ0Yx3JcIj/wNsJX1MxY="
FieldTypeRequiredDescription
typestringno"none" (default) or "password"
users[].namestringyesLogin name, matched against the client's user
users[].password_hashstringyesSCRAM-SHA-256 verifier (see below)

password mode uses SCRAM-SHA-256, the same mechanism as PostgreSQL's password_encryption = scram-sha-256. psql, JDBC, psycopg2/3, DBeaver, tokio-postgres and other modern clients support it. The password never crosses the network, and the config file holds only a salted verifier in PostgreSQL's pg_authid.rolpassword format, not the password. A verifier copied from PostgreSQL (SELECT rolpassword FROM pg_authid WHERE rolname = '…') works as-is. To generate one:

bash
echo -n 'my-password' | arneb hash-password     # or run it and type at the prompt
echo -n 'my-password' | arneb hash-password --iterations 100000

--iterations sets the PBKDF2 iteration count stored in the verifier. The default, 4096, matches PostgreSQL. For production, use a much higher count (for example 100000 or more) to slow down offline guessing if the config file leaks. The cost is paid by the client on every connection, so measure connection latency for your clients before going very high.

Behavior:

  • A wrong password or an unknown user is rejected with FATAL 28P01 password authentication failed for user "…". Unknown users get the same challenge as real ones, so an attacker can't tell which user names exist.
  • One startup handler serves both the Simple and Extended Query protocols, so every client path gets the same check.
  • The effective mode is logged at startup on the arneb::config target (pgwire authentication effective mode auth="password (scram-sha-256)" users=N).
  • Invalid [auth] config stops startup. This covers an unknown type, a malformed hash, duplicate or empty user names, type = "password" with no users, and a plaintext password key.

Limitations: [auth] is enforced only on the pgwire port. These listeners still accept anyone, so keep them on a trusted network:

  • the Trino client protocol port ([trino] port, default 8080, enabled by default): any X-Trino-User is accepted without a password. Startup logs a WARN on arneb::config when type = "password" is set and this listener is on. Where password protection is required, disable it with --no-trino or ARNEB_TRINO_ENABLED=false until the follow-up (HTTP Basic over TLS, checked against the same SCRAM verifiers) lands;
  • the Web UI (pgwire port + 1000);
  • the Flight RPC port between the coordinator and its workers.

The pgwire port does not support TLS yet. SCRAM keeps the password itself off the wire, but query traffic is still plaintext. Channel binding (SCRAM-SHA-256-PLUS) is not offered.

CLI Arguments ​

arneb [OPTIONS] [COMMAND]

Commands:
  hash-password      Read a password from stdin and print a SCRAM-SHA-256
                     verifier for [[auth.users]] password_hash

Options:
  --config <PATH>    Path to configuration file
  --port <PORT>      Override the pgwire port
  --role <ROLE>      Server role: standalone, coordinator, or worker

Example: Standalone with Local Files ​

toml
bind_address = "127.0.0.1"
port = 5432

[[tables]]
name = "lineitem"
path = "/data/tpch/lineitem.parquet"
format = "parquet"

[[tables]]
name = "orders"
path = "/data/tpch/orders.parquet"
format = "parquet"

Example: Distributed with Hive Catalog ​

toml
bind_address = "0.0.0.0"
port = 5432

[storage.s3]
region = "us-east-1"
endpoint = "http://minio:9000"
allow_http = true

[[catalogs]]
name = "datalake"
type = "hive"
metastore_uri = "hive-metastore:9083"
default_schema = "default"

[catalogs.storage.s3]
region = "us-east-1"
endpoint = "http://minio:9000"
allow_http = true