Compare commits

...
Author SHA1 Message Date
Codex Agent 714d64ed3c Merge remote-tracking branch 'origin/main' into p9/1733-persisted-keys 2026-08-31 03:03:14 +08:00
Codex Agent 137d4cba7f test(filemeta): pin persisted metadata key literals 2026-08-31 03:01:52 +08:00
唐小鸭andGitHub ec1cd606d3 fix(replication): surface object-lock denied purges and back off heal retries (#6900) 2026-08-30 18:59:55 +00:00
16af688a7a fix(rpc): reject unsigned v2 control mutations (#6905)
Co-authored-by: heihutu <[email protected]>
2026-08-30 18:17:43 +00:00
唐小鸭andGitHub 37b23a16da fix(replication): verify replica integrity and default to plain signed payloads (#6895) 2026-08-31 01:43:45 +08:00
006e9b7d28 test(e2e): cover four-node four-drive cluster topology (#6902)
Co-authored-by: heihutu <[email protected]>
2026-08-30 17:32:36 +00:00
d214c27583 perf(ecstore): consolidate non-inline read planning (#6892)
Co-authored-by: heihutu <[email protected]>
2026-08-30 17:15:48 +00:00
GatewayJandGitHub 8fd364a99c feat(s3-tables): support object-backed table rename (#6899) 2026-08-31 00:30:49 +08:00
c2d8488728 docs(architecture): reconcile generation contract (#6901)
Co-authored-by: heihutu <[email protected]>
2026-08-31 00:20:39 +08:00
唐小鸭andGitHub 1370434f3a fix(scanner): unblock quota usage baseline on never-converged sites (#6896) 2026-08-31 00:20:04 +08:00
唐小鸭andGitHub 5dde2c188c fix(replication): retry failed multipart aborts on bounded backoff (#6897) 2026-08-31 00:19:49 +08:00
2f9c75d04f perf(ecstore): reuse prepared metadata across pools (#6889)
Co-authored-by: heihutu <[email protected]>
2026-08-30 16:15:10 +00:00
唐小鸭andGitHub 9ee7b1221d fix(admin): replicate user secret-key rotation to peer sites (#6893) 2026-08-30 23:32:07 +08:00
Zhengchao AnandGitHub fcc3c7fb6b test(s3): promote passing compatibility cases (#6891) 2026-08-30 21:56:17 +08:00
Zhengchao AnandGitHub 01dc55ee5b docs(security): add unsigned presign header lesson (#6894) 2026-08-30 21:55:47 +08:00
3d24526704 fix(ecstore): preserve parity reserves for data-only GET (#6888)
fix(ecstore): hedge data-only GET with parity

Route the opt-in data-shards-only lockstep path through the bounded parity race and preserve deferred parity reserves across canceled hedges.

Co-authored-by: heihutu <[email protected]>
2026-08-30 20:16:32 +08:00
51532e19fb test(ecstore): cover multipart snapshot overwrite race (#6887)
test(ecstore): cover multipart GET overwrite snapshot

Co-authored-by: heihutu <[email protected]>
2026-08-30 12:15:02 +00:00
931ff60182 test(ci): refresh cluster nightly selection (#6886)
Co-authored-by: heihutu <[email protected]>
2026-08-30 12:10:07 +00:00
07212c4e26 perf(ecstore): gate quorum-aware GET early stop (#6885)
* perf(ecstore): add gated two-phase GET metadata reads

Co-Authored-By: heihutu <[email protected]>

* fix(ecstore): require data-shard coverage for read plans

Co-Authored-By: heihutu <[email protected]>

* perf(ecstore): avoid inline overhead in read plan rollout

Co-Authored-By: heihutu <[email protected]>

* perf(ecstore): accept quorum-complete read candidates

Co-Authored-By: heihutu <[email protected]>

---------

Co-authored-by: heihutu <[email protected]>
2026-08-30 09:33:35 +00:00
GatewayJandGitHub 4932af080b feat(s3select): expand typed JSON source paths (#6864) 2026-08-30 06:46:36 +00:00
GatewayJandGitHub d6f9a7c462 feat(table-catalog): vend credentials from LoadTable (#6878)
* feat(table-catalog): vend credentials from LoadTable

* fix(table-catalog): preserve entry-relative metadata paths
2026-08-30 06:33:28 +00:00
7345b49cf6 perf(ecstore): gate GET metadata timing when metrics off (#6879)
Co-authored-by: heihutu <[email protected]>
2026-08-30 05:43:08 +00:00
housemeandGitHub 4753e35035 chore(deps): update flake.lock (#6880) 2026-08-30 13:14:13 +08:00
GatewayJandGitHub 96239fc034 feat(s3select): report uncompressed input byte metrics (#6865) 2026-08-30 04:07:20 +00:00
hectorandGitHub b428875bed fix(ci): stabilize tier MQTT bootstrap on shared runner (#6877)
* fix(ci): isolate s3 compat temp file paths

* fix(ci): use rooted auto-testing s3 temp fix

* fix(ci): follow auto-testing main after temp-path merge

* fix(ci): stabilize tier mqtt bootstrap on shared runner
2026-08-30 11:14:28 +08:00
66 changed files with 8203 additions and 855 deletions
@@ -48,6 +48,7 @@ Update this file only when an advisory adds or changes a reusable lesson, affect
### S3 object actions, copy, multipart, and upload policy validation
- `GHSA-g8w9-qw9q-fghr`: a valid presigned `PutObject` accepted extra `x-amz-tagging`, website redirect, and storage-class headers omitted from `SignedHeaders`. Lesson: a presigned URL is a bounded capability; reject `x-amz-*` headers that are not cryptographically bound by the signature so unsigned metadata cannot change authorization, lifecycle, redirect, cost, or durability semantics.
- `GHSA-3ppv-fx5m-m749`: explicit `versionId` reads and copy sources authorized `s3:GetObject` instead of `s3:GetObjectVersion`. Lesson: version-specific object access must select version-specific actions for direct reads, `CopyObject`, and `UploadPartCopy`, with tests proving the backend is not reached on denial.
- `GHSA-x298-9x87-fvjq`: anonymous `ListObjectVersions` fell back to `ListBucket` and returned before public-access-block gates. Lesson: compatibility fallbacks must converge on the same post-authorization checks as direct grants, especially `RestrictPublicBuckets` and anonymous data-plane denies.
- `GHSA-mx42-j6wv-px98`: `UploadPartCopy` missed source authorization and allowed cross-bucket object exfiltration. Lesson: multipart copy must enforce the same source and destination contract as `CopyObject`.
@@ -119,7 +120,7 @@ Use these targeted searches when a diff touches security-sensitive code:
```bash
rg -n "validate_admin_request|check_permissions|AdminAction::|deny_only|is_allowed" rustfs crates
rg -n "authorize_operation|FtpsDriver|SftpDriver|RETR|MKD|SIZE|MDTM|CreateBucket|GetObject|HeadObject" crates/protocols rustfs
rg -n "UploadPartCopy|upload_part_copy|CompleteMultipart|PostObject|content-length-range|starts-with" rustfs crates
rg -n "UploadPartCopy|upload_part_copy|CompleteMultipart|PostObject|presign|SignedHeaders|content-length-range|starts-with" rustfs crates
rg -n "ListBucketVersions|GetObjectVersion|versionId|VersionId|ExistingObjectTag|ForAllValues|ForAnyValue|POLICY_PLUGIN|opa" rustfs crates
rg -n "normalize_extract_entry_key|Snowball|auto-extract|PathBuf::join|canonicalize|\\.\\.|x-forwarded-for|x-real-ip|SourceIp" rustfs crates
rg -n "DEFAULT_SECRET|DEFAULT_ACCESS|TEST_PRIVATE_KEY|rustfs rpc|RUSTFS_RPC_SECRET" rustfs crates
@@ -136,6 +137,7 @@ rg -n "deny_unknown_fields|serde.default|as u32|as usize|as i32" rustfs crates
- Protocol frontend authz fixes: include denied `RETR`, `SIZE`/`MDTM`, `MKD`, bucket probe, and sibling allowed-operation cases, and assert denied paths do not reach the storage backend.
- IAM fixes: include import/update/list service-account cases with attacker-controlled parent, claims, access key, secret key, and policy.
- Copy/upload fixes: include cross-bucket, cross-user, source-denied, destination-denied, copy-source-condition, and multipart completion cases.
- Presigned upload fixes: include a valid presign with extra unsigned tagging, redirect, and storage-class headers; require rejection before storage access, and verify explicitly signed equivalents still work.
- Version-action fixes: include historical UUID, explicit current version, `null`, range, partNumber, presigned, STS/session, service-account, anonymous bucket-policy, copy source, and multipart-copy source cases.
- Policy-condition fixes: include reserved-key header collisions, missing keys, partially overlapping multi-value sets, plugin mode, and built-in policy mode.
- Path fixes: include encoded traversal, absolute path, nested traversal, archive entries with `..`, valid object keys that resemble traversal text but should be rejected, and canonical bucket/prefix boundary checks.
+1 -1
View File
@@ -1 +1 @@
sha256=9b9bc336b43b70d0e06e0adb5455bf035bb18945d85d60936eb6fe4d48e0e680
sha256=26003ce03eca11391d1c080491e4f408526717e4b47967b62db08abe1edd189a
+26 -6
View File
@@ -60,6 +60,8 @@ jobs:
- name: Cleanup environment (before)
run: |
set -euo pipefail
sudo docker rm -f rustfs-test-mqtt >/dev/null 2>&1 || true
sudo rm -f /tmp/rustfs-mosquitto.conf
read -r -a NODES <<< "${RUSTFS_NODES:-vm000 vm001 vm002}"
SSH_USER="${RUSTFS_SSH_USER:-azureuser}"
for node in "${NODES[@]}"; do
@@ -77,15 +79,31 @@ jobs:
- name: Ensure MQTT broker + clients
run: |
set -euo pipefail
if ! command -v mosquitto_sub >/dev/null 2>&1; then
sudo apt-get update
sudo apt-get install -y mosquitto mosquitto-clients
sudo apt-get install -y mosquitto-clients
fi
sudo mkdir -p /etc/mosquitto/conf.d
printf 'listener 1883 0.0.0.0\nallow_anonymous true\n' | sudo tee /etc/mosquitto/conf.d/rustfs-test.conf >/dev/null
sudo systemctl restart mosquitto
sleep 2
ss -tln 2>/dev/null | grep -q ':1883' || { echo 'mosquitto not listening on 1883'; exit 1; }
command -v docker >/dev/null 2>&1 || { echo 'docker not found on runner'; exit 1; }
sudo docker rm -f rustfs-test-mqtt >/dev/null 2>&1 || true
cat <<'EOF' | sudo tee /tmp/rustfs-mosquitto.conf >/dev/null
listener 1883 0.0.0.0
allow_anonymous true
EOF
sudo docker run -d --name rustfs-test-mqtt -p 1883:1883 \
-v /tmp/rustfs-mosquitto.conf:/mosquitto/config/mosquitto.conf:ro \
eclipse-mosquitto:2 >/dev/null
for _ in {1..10}; do
if ss -tln 2>/dev/null | grep -q ':1883'; then
break
fi
sleep 1
done
ss -tln 2>/dev/null | grep -q ':1883' || {
echo 'mosquitto container is not listening on 1883'
sudo docker logs rustfs-test-mqtt || true
exit 1
}
- name: Run tier suite
id: test
@@ -150,6 +168,8 @@ jobs:
if: always()
run: |
set -euo pipefail
sudo docker rm -f rustfs-test-mqtt >/dev/null 2>&1 || true
sudo rm -f /tmp/rustfs-mosquitto.conf
read -r -a NODES <<< "${RUSTFS_NODES:-vm000 vm001 vm002}"
SSH_USER="${RUSTFS_SSH_USER:-azureuser}"
for node in "${NODES[@]}"; do
@@ -76,6 +76,28 @@ async fn cluster_multidrive_single_pool_smoke() -> TestResult {
Ok(())
}
/// 4 nodes x 4 drives, single pool: exercise the maximum local erasure layout
/// supported by the cluster harness. This remains in the nightly lane because
/// it starts four real server processes and sixteen data directories.
#[tokio::test]
async fn cluster_four_node_four_drive_single_pool_smoke() -> TestResult {
crate::common::init_logging();
let mut cluster = RustFSTestClusterEnvironment::with_topology(ClusterTopology::single_pool_multidrive(4, 4)).await?;
let volumes = cluster.rustfs_volumes_arg();
assert_eq!(volumes.split(' ').count(), 16, "expected 16 explicit endpoints, got: {volumes}");
assert!(!volumes.contains('{'), "single-pool layout must not use ellipses: {volumes}");
assert!(cluster.nodes.iter().all(|node| node.data_dirs.len() == 4));
cluster.start().await?;
cluster.create_test_bucket(BUCKET).await?;
let payload = vec![0x3Cu8; 1024 * 1024];
put_get_roundtrip(&cluster, "multidrive-4/object", &payload).await?;
Ok(())
}
/// Two single-node pools, 2 drives each: the multi-pool layout boots and
/// round-trips. Every pool is a distinct erasure pool (`pool_idx` 0 and 1).
#[tokio::test]
+287 -1
View File
@@ -17,7 +17,8 @@ use crate::common::{RustFSTestEnvironment, init_logging};
use aws_sdk_s3::Client;
use aws_sdk_s3::error::ProvideErrorMetadata;
use aws_sdk_s3::types::{
CsvInput, CsvOutput, ExpressionType, FileHeaderInfo, InputSerialization, JsonInput, JsonOutput, JsonType, OutputSerialization,
CsvInput, CsvOutput, ExpressionType, FileHeaderInfo, InputSerialization, JsonInput, JsonOutput, JsonType,
OutputSerialization, RequestProgress,
};
use bytes::Bytes;
use std::error::Error;
@@ -26,6 +27,9 @@ use std::time::Duration;
const BUCKET: &str = "test-sql-bucket";
const CSV_OBJECT: &str = "test-data.csv";
const JSON_OBJECT: &str = "test-data.json";
const JSON_DOCUMENT_OBJECT: &str = "nested-data.json";
const JSON_ROOT_ARRAY_OBJECT: &str = "root-array.json";
const JSON_ROOT_SCALAR_ARRAY_OBJECT: &str = "root-scalars.json";
const SELECT_RESPONSE_TIMEOUT: Duration = Duration::from_secs(30);
type TestResult<T> = Result<T, Box<dyn Error + Send + Sync>>;
@@ -73,6 +77,51 @@ async fn upload_test_json(client: &Client) -> TestResult<()> {
Ok(())
}
async fn upload_nested_json_document(client: &Client) -> TestResult<()> {
let json_data = r#"{"departments":[{"employees":[{"name":"Alice","active":true},{"name":"Bob","active":false}]},{"employees":[{"name":"Charlie","active":true}]}]}"#;
client
.put_object()
.bucket(BUCKET)
.key(JSON_DOCUMENT_OBJECT)
.body(Bytes::from_static(json_data.as_bytes()).into())
.send()
.await?;
client
.put_object()
.bucket(BUCKET)
.key(JSON_ROOT_ARRAY_OBJECT)
.body(Bytes::from_static(br#"[{"name":"Alice"},{"name":"Bob"}]"#).into())
.send()
.await?;
client
.put_object()
.bucket(BUCKET)
.key(JSON_ROOT_SCALAR_ARRAY_OBJECT)
.body(Bytes::from_static(b"[1,2]").into())
.send()
.await?;
Ok(())
}
async fn select_json_document(client: &Client, key: &str, expression: &str) -> TestResult<String> {
let response = client
.select_object_content()
.bucket(BUCKET)
.key(key)
.expression(expression)
.expression_type(ExpressionType::Sql)
.input_serialization(
InputSerialization::builder()
.json(JsonInput::builder().set_type(Some(JsonType::Document)).build())
.build(),
)
.output_serialization(OutputSerialization::builder().json(JsonOutput::builder().build()).build())
.send()
.await?;
process_select_response(response).await
}
async fn process_select_response(
mut event_stream: aws_sdk_s3::operation::select_object_content::SelectObjectContentOutput,
) -> TestResult<String> {
@@ -104,6 +153,142 @@ async fn process_select_response(
.map_err(|_| -> Box<dyn Error + Send + Sync> { "Select response timed out".into() })?
}
async fn assert_input_byte_stats(
client: &Client,
object: &str,
body: &[u8],
expression: &str,
input_serialization: InputSerialization,
output_serialization: OutputSerialization,
progress_enabled: bool,
) -> TestResult<()> {
client
.put_object()
.bucket(BUCKET)
.key(object)
.body(Bytes::copy_from_slice(body).into())
.send()
.await?;
let mut request = client
.select_object_content()
.bucket(BUCKET)
.key(object)
.expression(expression)
.expression_type(ExpressionType::Sql)
.input_serialization(input_serialization)
.output_serialization(output_serialization);
if progress_enabled {
request = request.request_progress(RequestProgress::builder().enabled(true).build());
}
let response = request.send().await?;
let mut payload = response.payload;
let mut records_len = 0_u64;
let mut last_progress: Option<aws_sdk_s3::types::Progress> = None;
let mut stats = None;
let mut saw_end = false;
while let Some(event) = payload.recv().await? {
match event {
aws_sdk_s3::types::SelectObjectContentEventStream::Records(records) => {
if let Some(bytes) = records.payload {
records_len = records_len.saturating_add(u64::try_from(bytes.as_ref().len())?);
}
}
aws_sdk_s3::types::SelectObjectContentEventStream::Progress(event) => {
let details = event.details.ok_or("Progress event did not contain details")?;
if let Some(previous) = last_progress.as_ref() {
assert!(details.bytes_scanned() >= previous.bytes_scanned());
assert!(details.bytes_processed() >= previous.bytes_processed());
assert!(details.bytes_returned() >= previous.bytes_returned());
}
last_progress = Some(details);
}
aws_sdk_s3::types::SelectObjectContentEventStream::Stats(event) => stats = event.details,
aws_sdk_s3::types::SelectObjectContentEventStream::End(_) => {
saw_end = true;
break;
}
_ => {}
}
}
let stats = stats.ok_or("Select response ended without a Stats event")?;
let input_len = i64::try_from(body.len())?;
assert_eq!(stats.bytes_scanned(), Some(input_len));
assert_eq!(stats.bytes_processed(), Some(input_len));
assert_eq!(stats.bytes_returned(), Some(i64::try_from(records_len)?));
if progress_enabled {
let progress = last_progress.ok_or("Select response ended without a Progress event")?;
assert_eq!(progress.bytes_scanned(), stats.bytes_scanned());
assert_eq!(progress.bytes_processed(), stats.bytes_processed());
assert_eq!(progress.bytes_returned(), stats.bytes_returned());
} else {
assert!(last_progress.is_none(), "disabled request progress emitted a Progress event");
}
assert!(saw_end, "Select response ended without an End event");
Ok(())
}
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
async fn test_select_object_content_reports_input_byte_stats() -> TestResult<()> {
const CSV_BODY: &[u8] = b"name,age\nAlice,30\nBob,25\n";
const JSON_LINES_BODY: &[u8] = b"{\"name\":\"Alice\"}\n{\"name\":\"Bob\"}\n";
const JSON_DOCUMENT_BODY: &[u8] = b"[{\"name\":\"Alice\"},{\"name\":\"Bob\"}]";
let (_env, client) = create_test_environment().await?;
setup_test_bucket(&client).await?;
assert_input_byte_stats(
&client,
"input-metrics.csv",
CSV_BODY,
"SELECT name FROM S3Object",
InputSerialization::builder()
.csv(CsvInput::builder().file_header_info(FileHeaderInfo::Use).build())
.build(),
OutputSerialization::builder().csv(CsvOutput::builder().build()).build(),
true,
)
.await?;
assert_input_byte_stats(
&client,
"input-metrics.jsonl",
JSON_LINES_BODY,
"SELECT name FROM S3Object",
InputSerialization::builder()
.json(JsonInput::builder().set_type(Some(JsonType::Lines)).build())
.build(),
OutputSerialization::builder().json(JsonOutput::builder().build()).build(),
true,
)
.await?;
assert_input_byte_stats(
&client,
"input-metrics.json",
JSON_DOCUMENT_BODY,
"SELECT name FROM S3Object",
InputSerialization::builder()
.json(JsonInput::builder().set_type(Some(JsonType::Document)).build())
.build(),
OutputSerialization::builder().json(JsonOutput::builder().build()).build(),
true,
)
.await?;
assert_input_byte_stats(
&client,
"input-metrics-without-progress.csv",
CSV_BODY,
"SELECT name FROM S3Object",
InputSerialization::builder()
.csv(CsvInput::builder().file_header_info(FileHeaderInfo::Use).build())
.build(),
OutputSerialization::builder().csv(CsvOutput::builder().build()).build(),
false,
)
.await?;
Ok(())
}
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
async fn test_select_object_content_csv_basic() -> TestResult<()> {
let (_env, client) = create_test_environment().await?;
@@ -228,6 +413,107 @@ async fn test_select_object_content_json_basic() -> TestResult<()> {
Ok(())
}
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
async fn test_select_object_content_nested_json_source_path() -> TestResult<()> {
let (_env, client) = create_test_environment().await?;
setup_test_bucket(&client).await?;
upload_nested_json_document(&client).await?;
let result = select_json_document(
&client,
JSON_DOCUMENT_OBJECT,
"SELECT e.name FROM S3Object[*].departments[*].employees[*] AS e WHERE e.active = true",
)
.await?;
let names: Vec<String> = result
.lines()
.filter(|line| !line.trim().is_empty())
.map(|line| -> TestResult<String> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["name"].as_str().ok_or("missing name field")?.to_string())
})
.collect::<TestResult<_>>()?;
assert_eq!(names, vec!["Alice", "Charlie"]);
let terminal_scalars = select_json_document(
&client,
JSON_DOCUMENT_OBJECT,
"SELECT NAME FROM S3Object[*].DEPARTMENTS[*].employees[*].NAME",
)
.await?;
let scalar_names: Vec<String> = terminal_scalars
.lines()
.map(|line| -> TestResult<String> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["name"].as_str().ok_or("missing scalar name field")?.to_string())
})
.collect::<TestResult<_>>()?;
assert_eq!(scalar_names, vec!["Alice", "Bob", "Charlie"]);
let aliased_scalars = select_json_document(
&client,
JSON_DOCUMENT_OBJECT,
"SELECT v FROM S3Object[*].departments[*].employees[*].name AS v",
)
.await?;
let aliased_names: Vec<String> = aliased_scalars
.lines()
.map(|line| -> TestResult<String> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["v"].as_str().ok_or("missing aliased scalar field")?.to_string())
})
.collect::<TestResult<_>>()?;
assert_eq!(aliased_names, vec!["Alice", "Bob", "Charlie"]);
let root_array = select_json_document(&client, JSON_ROOT_ARRAY_OBJECT, "SELECT c.name FROM S3Object[*][*] AS c").await?;
let root_names: Vec<String> = root_array
.lines()
.map(|line| -> TestResult<String> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["name"].as_str().ok_or("missing root-array name field")?.to_string())
})
.collect::<TestResult<_>>()?;
assert_eq!(root_names, vec!["Alice", "Bob"]);
let root_index = select_json_document(&client, JSON_ROOT_ARRAY_OBJECT, "SELECT c.name FROM S3Object[*][0] AS c").await?;
let root_index_value: serde_json::Value = serde_json::from_str(root_index.trim())?;
assert_eq!(root_index_value["name"], "Alice");
let root_scalars = select_json_document(&client, JSON_ROOT_SCALAR_ARRAY_OBJECT, "SELECT V FROM S3Object AS V").await?;
let scalar_values: Vec<i64> = root_scalars
.lines()
.map(|line| -> TestResult<i64> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["v"].as_i64().ok_or("missing root scalar value")?)
})
.collect::<TestResult<_>>()?;
assert_eq!(scalar_values, vec![1, 2]);
let implicit_root_scalars =
select_json_document(&client, JSON_ROOT_SCALAR_ARRAY_OBJECT, "SELECT S3Object FROM S3Object").await?;
let implicit_scalar_values: Vec<i64> = implicit_root_scalars
.lines()
.map(|line| -> TestResult<i64> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["s3object"].as_i64().ok_or("missing implicit root scalar value")?)
})
.collect::<TestResult<_>>()?;
assert_eq!(implicit_scalar_values, vec![1, 2]);
let quoted_root_scalars =
select_json_document(&client, JSON_ROOT_SCALAR_ARRAY_OBJECT, "SELECT \"S3Object\" FROM \"S3Object\"").await?;
let quoted_scalar_values: Vec<i64> = quoted_root_scalars
.lines()
.map(|line| -> TestResult<i64> {
let value: serde_json::Value = serde_json::from_str(line)?;
Ok(value["S3Object"].as_i64().ok_or("missing quoted root scalar value")?)
})
.collect::<TestResult<_>>()?;
assert_eq!(quoted_scalar_values, vec![1, 2]);
Ok(())
}
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
async fn test_select_object_content_csv_limit() -> TestResult<()> {
let (_env, client) = create_test_environment().await?;
+2 -2
View File
@@ -461,8 +461,8 @@ pub mod rpc {
tonic_boot_epoch_challenge, tonic_boot_epoch_response_headers, tonic_rpc_auth_failure_reason,
verify_ns_scanner_capability, verify_ns_scanner_capability_with_tier_registry_generation, verify_put_file_auth_trailer,
verify_put_file_capability, verify_rpc_signature, verify_tonic_boot_epoch_response, verify_tonic_canonical_body_digest,
verify_tonic_mutation_body_digest, verify_tonic_rpc_response_proof, verify_tonic_rpc_signature,
verify_tonic_rpc_signature_with_bootstrap,
verify_tonic_mutation_body_digest, verify_tonic_mutation_body_digest_reject_unsigned, verify_tonic_rpc_response_proof,
verify_tonic_rpc_signature, verify_tonic_rpc_signature_with_bootstrap,
};
}
+200 -5
View File
@@ -24,6 +24,7 @@ use crate::runtime::sources as runtime_sources;
use aws_credential_types::Credentials as SdkCredentials;
use aws_credential_types::provider::{ProvideCredentials, error::CredentialsError, future};
use aws_sdk_s3::config::Region as SdkRegion;
use aws_sdk_s3::config::RequestChecksumCalculation;
use aws_sdk_s3::config::SharedHttpClient;
use aws_sdk_s3::error::ProvideErrorMetadata;
use aws_sdk_s3::error::SdkError;
@@ -39,6 +40,7 @@ use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::types::Tagging as SdkTagging;
use aws_sdk_s3::types::{
ChecksumMode, CompletedMultipartUpload, CompletedPart, ObjectLockLegalHoldStatus, ObjectLockRetentionMode,
ServerSideEncryption,
};
use aws_sdk_s3::{Client as S3Client, Config as S3Config, operation::head_object::HeadObjectOutput};
use aws_sdk_s3::{config::SharedCredentialsProvider, types::BucketVersioningStatus};
@@ -1071,7 +1073,8 @@ impl BucketTargetSys {
.endpoint_url(endpoint.clone())
.credentials_provider(SharedCredentialsProvider::new(RemoteTargetCredentialsProvider { credentials: creds }))
.region(SdkRegion::new(target.region.clone()))
.behavior_version(aws_sdk_s3::config::BehaviorVersion::latest());
.behavior_version(aws_sdk_s3::config::BehaviorVersion::latest())
.request_checksum_calculation(replication_request_checksum_calculation());
if should_force_path_style(target) {
config_builder = config_builder.force_path_style(true);
@@ -1367,6 +1370,25 @@ fn loopback_replication_targets_allowed() -> bool {
.unwrap_or(false)
}
const REPLICATION_STREAMING_CHECKSUMS_ENV: &str = "RUSTFS_REPLICATION_STREAMING_CHECKSUMS";
/// Streaming trailer checksums make the SDK frame request bodies as
/// `aws-chunked`; a target that does not decode that framing stores the frames
/// verbatim, silently corrupting every replica while the transfer itself
/// succeeds (#6853). Plain signed payloads are the compatible default; the env
/// knob restores trailer checksums for fleets whose targets are all known to
/// decode them.
fn replication_request_checksum_calculation() -> RequestChecksumCalculation {
if std::env::var(REPLICATION_STREAMING_CHECKSUMS_ENV)
.map(|v| v.eq_ignore_ascii_case("true") || v == "1")
.unwrap_or(false)
{
RequestChecksumCalculation::WhenSupported
} else {
RequestChecksumCalculation::WhenRequired
}
}
fn validate_replication_target_endpoint(url: &Url) -> Result<(), OutboundUrlError> {
validate_replication_target_endpoint_inner(url, loopback_replication_targets_allowed())
}
@@ -1746,6 +1768,17 @@ impl Default for AdvancedPutOptions {
}
}
/// The subset of the target's PutObject response replication audits.
#[derive(Debug, Clone)]
pub struct RemotePutObjectResponse {
/// Version id the target assigned (`x-amz-version-id`).
pub version_id: Option<String>,
/// ETag of what the target stored; `None` when the target withheld it or
/// when its encryption mode (SSE-KMS / SSE-C) makes it incomparable to
/// the source ETag. `None` is therefore "not decidable", never evidence.
pub etag: Option<String>,
}
#[derive(Clone)]
pub struct PutObjectOptions {
pub user_metadata: HashMap<String, String>,
@@ -2291,7 +2324,9 @@ impl TargetClient {
/// On success returns the version id the target assigned (from
/// `x-amz-version-id`), letting callers audit the version-identity
/// contract — a target that adopts the source version echoes it back.
/// contract — a target that adopts the source version echoes it back
/// together with the ETag of what the target actually stored, so callers
/// can detect a target that persisted transformed bytes (#6853).
pub async fn put_object(
&self,
bucket: &str,
@@ -2299,7 +2334,7 @@ impl TargetClient {
size: i64,
body: ByteStream,
opts: &PutObjectOptions,
) -> Result<Option<String>, S3ClientError> {
) -> Result<RemotePutObjectResponse, S3ClientError> {
let mut headers = opts.header();
let builder = self.client.put_object();
@@ -2334,7 +2369,25 @@ impl TargetClient {
.send()
.await
{
Ok(output) => Ok(output.version_id().map(ToOwned::to_owned)),
Ok(output) => {
// Under SSE-KMS/DSSE or SSE-C the target's ETag is not the MD5
// of the stored plaintext, so it cannot be compared against the
// source ETag; withhold it rather than let a caller conclude
// corruption from an opaque value.
let etag_comparable = output.sse_customer_algorithm().is_none()
&& !matches!(
output.server_side_encryption(),
Some(ServerSideEncryption::AwsKms) | Some(ServerSideEncryption::AwsKmsDsse)
);
Ok(RemotePutObjectResponse {
version_id: output.version_id().map(ToOwned::to_owned),
etag: if etag_comparable {
output.e_tag().map(ToOwned::to_owned)
} else {
None
},
})
}
Err(e) => match e {
SdkError::ServiceError(service_err) => {
let err = service_err.into_err();
@@ -2673,6 +2726,145 @@ mod tests {
}
}
type RecordedHeaders = Arc<std::sync::Mutex<Vec<Vec<(String, String)>>>>;
/// Records full request headers and answers with canned response headers,
/// for asserting wire framing and response parsing.
#[derive(Clone, Debug)]
struct RecordingHeaderConnector {
request_headers: RecordedHeaders,
response_headers: Vec<(String, String)>,
}
impl SmithyHttpConnector for RecordingHeaderConnector {
fn call(&self, request: HttpRequest) -> HttpConnectorFuture {
self.request_headers
.lock()
.expect("recorded header lock should not be poisoned")
.push(
request
.headers()
.iter()
.map(|(k, v)| (k.to_string(), v.to_string()))
.collect(),
);
let mut response = HttpResponse::new(
aws_smithy_runtime_api::http::StatusCode::try_from(200_u16).expect("200 should be a valid response status"),
SdkBody::empty(),
);
for (name, value) in &self.response_headers {
response.headers_mut().insert(name.clone(), value.clone());
}
HttpConnectorFuture::ready(Ok(response))
}
}
fn header_recording_target_client(response_headers: Vec<(String, String)>) -> (TargetClient, RecordedHeaders) {
let request_headers: RecordedHeaders = Arc::new(std::sync::Mutex::new(Vec::new()));
let connector = SharedHttpConnector::new(RecordingHeaderConnector {
request_headers: Arc::clone(&request_headers),
response_headers,
});
let http_client = http_client_fn(move |_settings, _components| connector.clone());
let client = s3_client_for_test(443, Some(http_client));
(
TargetClient {
endpoint: "https://localhost:443".to_string(),
credentials: None,
bucket: "target-bucket".to_string(),
storage_class: String::new(),
disable_proxy: false,
arn: "arn:rustfs:replication:us-east-1:target:bucket".to_string(),
reset_id: String::new(),
secure: true,
health_check_duration: Duration::from_secs(5),
replicate_sync: false,
client: Arc::new(client),
},
request_headers,
)
}
fn streaming_test_body(payload: &'static [u8]) -> ByteStream {
let stream = tokio_util::io::ReaderStream::new(std::io::Cursor::new(payload));
let body = http_body_util::StreamBody::new(futures::StreamExt::map(stream, |r| r.map(http_body::Frame::data)));
ByteStream::new(SdkBody::from_body_1_x(body))
}
#[test]
fn replication_checksums_default_to_plain_payloads() {
assert!(matches!(
replication_request_checksum_calculation(),
RequestChecksumCalculation::WhenRequired
));
}
#[tokio::test]
async fn replication_put_object_sends_plain_signed_payloads_by_default() {
let (client, recorded) = header_recording_target_client(Vec::new());
client
.put_object("target-bucket", "object", 4, streaming_test_body(b"data"), &PutObjectOptions::default())
.await
.expect("recorded put_object should succeed");
let recorded = recorded.lock().expect("recorded header lock should not be poisoned");
let headers = &recorded[0];
let header = |name: &str| {
headers
.iter()
.find(|(k, _)| k.eq_ignore_ascii_case(name))
.map(|(_, v)| v.as_str())
};
// The #6853 regression shape: trailer checksums force aws-chunked
// framing, which a non-decoding target stores verbatim as the object.
assert_eq!(header("x-amz-trailer"), None, "streaming uploads must not carry a trailer checksum");
assert!(
header("content-encoding").is_none_or(|v| !v.contains("aws-chunked")),
"streaming uploads must not be aws-chunked framed"
);
assert_eq!(header("x-amz-decoded-content-length"), None);
assert_eq!(header("content-length"), Some("4"));
}
#[tokio::test]
async fn put_object_returns_the_etag_the_target_stored() {
let (client, _) =
header_recording_target_client(vec![("etag".to_string(), "\"9a0364b9e99bb480dd25e1f0284c8555\"".to_string())]);
let response = client
.put_object(
"target-bucket",
"object",
4,
ByteStream::from_static(b"data"),
&PutObjectOptions::default(),
)
.await
.expect("recorded put_object should succeed");
assert_eq!(response.etag.as_deref(), Some("\"9a0364b9e99bb480dd25e1f0284c8555\""));
}
#[tokio::test]
async fn put_object_withholds_the_etag_under_target_side_kms() {
let (client, _) = header_recording_target_client(vec![
("etag".to_string(), "\"9a0364b9e99bb480dd25e1f0284c8555\"".to_string()),
("x-amz-server-side-encryption".to_string(), "aws:kms".to_string()),
]);
let response = client
.put_object(
"target-bucket",
"object",
4,
ByteStream::from_static(b"data"),
&PutObjectOptions::default(),
)
.await
.expect("recorded put_object should succeed");
assert!(
response.etag.is_none(),
"a KMS-encrypted replica's etag is not the content MD5 and must be withheld"
);
}
#[derive(Clone, Debug)]
struct RecordingAuthConnector {
signed_requests: Arc<std::sync::Mutex<Vec<(bool, bool)>>>,
@@ -2969,7 +3161,10 @@ mod tests {
.credentials_provider(SharedCredentialsProvider::new(credentials))
.region(SdkRegion::new("us-east-1"))
.force_path_style(true)
.behavior_version(aws_sdk_s3::config::BehaviorVersion::latest());
.behavior_version(aws_sdk_s3::config::BehaviorVersion::latest())
// Mirror the production remote-target builder so recorded requests
// exercise the same checksum/framing behavior (#6853).
.request_checksum_calculation(replication_request_checksum_calculation());
if let Some(http_client) = http_client {
config = config.http_client(http_client);
}
@@ -20,8 +20,8 @@ pub use rustfs_replication::{
pub(crate) use rustfs_replication::{
ReplicationDeleteSource, ReplicationMultipartPartInput, ReplicationResyncTargetObject, delete_marker_purge_mrf_entry,
delete_marker_purge_version_id, delete_replication_creates_marker, delete_replication_missing_source_decision,
delete_replication_object_opts, heal_uses_delete_replication_path, is_retryable_delete_replication_head_error,
is_version_delete_replication, replicate_delete_outcome, replication_etags_match, replication_multipart_complete_actual_size,
replication_multipart_part_plan, resync_existing_delete_replication_info, resync_target_for_object,
should_retry_delete_marker_purge, target_delete_version_id,
delete_replication_object_opts, heal_uses_delete_replication_path, is_object_lock_denied_delete,
is_retryable_delete_replication_head_error, is_version_delete_replication, replicate_delete_outcome, replication_etags_match,
replication_multipart_complete_actual_size, replication_multipart_part_plan, resync_existing_delete_replication_info,
resync_target_for_object, should_retry_delete_marker_purge, single_part_replica_etag_mismatch, target_delete_version_id,
};
@@ -3177,6 +3177,19 @@ pub(crate) async fn queue_replication_heal_internal(
}
}
ReplicationHealQueueAction::QueueDelete(dv) => {
// A purge the peer denied under object lock cannot succeed until
// the lock lapses (#6850); requeuing it every heal cycle only
// burns bandwidth and failure counters. The backoff expires on
// its own, so the purge is probed again — and converges — once
// the retention window has a chance of being over.
if super::replication_object_decision_boundary::is_version_delete_replication(&dv.delete_object)
&& super::replication_resyncer::object_lock_denied_purge_backoff_active(&dv)
{
return ReplicationHealQueueResult {
object_info: roi,
admission: ReplicationQueueAdmission::Skipped,
};
}
let admission = if let Some(pool) = runtime_sources::replication_pool() {
pool.queue_replica_delete_task(dv).await
} else {
@@ -30,10 +30,10 @@ use super::replication_msgp_boundary::ReplicationMsgpCodec;
use super::replication_object_config::{ReplicationConfig, get_replication_config, must_replicate};
use super::replication_object_decision_boundary::{
MustReplicateOptions, ReplicationMultipartPartInput, delete_marker_purge_mrf_entry, delete_marker_purge_version_id,
delete_replication_creates_marker, heal_uses_delete_replication_path, is_retryable_delete_replication_head_error,
is_version_delete_replication, replicate_delete_outcome, replication_etags_match, replication_multipart_complete_actual_size,
replication_multipart_part_plan, resync_existing_delete_replication_info, should_retry_delete_marker_purge,
target_delete_version_id,
delete_replication_creates_marker, heal_uses_delete_replication_path, is_object_lock_denied_delete,
is_retryable_delete_replication_head_error, is_version_delete_replication, replicate_delete_outcome, replication_etags_match,
replication_multipart_complete_actual_size, replication_multipart_part_plan, resync_existing_delete_replication_info,
should_retry_delete_marker_purge, single_part_replica_etag_mismatch, target_delete_version_id,
};
use super::replication_queue_boundary::{DeletedObjectReplicationInfo, ReplicationQueueAdmission};
use super::replication_resync_boundary::ResyncStatusType;
@@ -54,7 +54,7 @@ use super::replication_storage_boundary::{
};
use super::replication_target_boundary::{
ERR_REPLICATION_SSEC_PASSTHROUGH_UNSUPPORTED, HeadObjectSdkError, PutObjectOptions, PutObjectPartOptions,
ReplicationTargetStore, S3ClientError, SsecPassthroughCapability, SsecPassthroughGate, TargetClient,
RemotePutObjectResponse, ReplicationTargetStore, S3ClientError, SsecPassthroughCapability, SsecPassthroughGate, TargetClient,
is_replication_target_offline_error, replication_action_for_target_head, replication_complete_multipart_options,
replication_delete_marker_purge_remove_options, replication_delete_remove_options, replication_force_delete_remove_options,
replication_object_is_ssec_encrypted, replication_put_object_header_size, replication_put_object_options,
@@ -96,7 +96,7 @@ use tokio::task::{JoinHandle, JoinSet};
use tokio::time::Duration as TokioDuration;
use tokio_util::io::ReaderStream;
use tokio_util::sync::CancellationToken;
use tracing::{debug, error, instrument, trace, warn};
use tracing::{debug, error, info, instrument, trace, warn};
const BACKGROUND_WALKDIR_TIMEOUT: TokioDuration = TokioDuration::from_secs(60);
const ENV_REPL_RESYNC_MAX_JOBS: &str = "RUSTFS_REPL_RESYNC_MAX_JOBS";
@@ -112,11 +112,13 @@ const EVENT_REPLICATION_DELETE_SKIPPED: &str = "replication_delete_skipped";
const EVENT_REPLICATION_FORCE_DELETE_SKIPPED: &str = "replication_force_delete_skipped";
const EVENT_RESYNC_TASK_FAILED: &str = "replication_resync_task_failed";
const EVENT_RESYNC_TARGET_OPERATION_FAILED: &str = "replication_resync_target_operation_failed";
const EVENT_REPLICATION_ABORT_RETRY_RESOLVED: &str = "replication_abort_retry_resolved";
const EVENT_RESYNC_RUNTIME_CHANNEL_FAILED: &str = "replication_resync_runtime_channel_failed";
const EVENT_DELETE_MARKER_PURGE_FAILED: &str = "replication_delete_marker_purge_failed";
const EVENT_DELETE_MARKER_PURGE_MRF: &str = "replication_delete_marker_purge_mrf";
const METRIC_DELETE_MARKER_PURGE_TOTAL: &str = "rustfs_replication_delete_marker_purge_total";
const EVENT_REPLICATION_VERSION_IDENTITY_DRIFT: &str = "replication_version_identity_drift";
const EVENT_REPLICATION_PURGE_OBJECT_LOCK_DENIED: &str = "replication_purge_object_lock_denied";
#[allow(
dead_code,
@@ -194,6 +196,123 @@ const METRIC_VERSION_IDENTITY_DRIFT_TOTAL: &str = "rustfs_replication_version_id
/// after a restart is acceptable.
static VERSION_IDENTITY_WARNED_ARNS: LazyLock<StdMutex<HashSet<String>>> = LazyLock::new(|| StdMutex::new(HashSet::new()));
/// Version purges the peer denied under object lock (#6850). Replication
/// carries no governance bypass, so such a purge cannot succeed until the
/// lock on the replica lapses — retrying every heal cycle only burns
/// bandwidth and failure counters. Entries suppress heal requeues for the
/// backoff window; after it expires one probe runs again, so the purge still
/// converges on its own once retention ends. In-process only: a restart
/// costs at most one extra probe per entry.
const OBJECT_LOCK_DENIED_PURGE_BACKOFF: std::time::Duration = std::time::Duration::from_secs(60 * 60);
const OBJECT_LOCK_DENIED_PURGE_CACHE_MAX: usize = 4096;
type ObjectLockDeniedPurgeKey = (String, String, String);
struct ObjectLockDeniedPurge {
denied_at: std::time::Instant,
denied_arns: HashSet<String>,
}
static OBJECT_LOCK_DENIED_PURGES: LazyLock<StdMutex<HashMap<ObjectLockDeniedPurgeKey, ObjectLockDeniedPurge>>> =
LazyLock::new(|| StdMutex::new(HashMap::new()));
fn object_lock_denied_purge_key(dobj: &DeletedObjectReplicationInfo) -> ObjectLockDeniedPurgeKey {
let version_id = dobj
.delete_object
.delete_marker_version_id
.or(dobj.delete_object.version_id)
.unwrap_or_default();
(dobj.bucket.clone(), dobj.delete_object.object_name.clone(), version_id.to_string())
}
fn record_object_lock_denied_purge(dobj: &DeletedObjectReplicationInfo, arn: &str) {
let mut denied = OBJECT_LOCK_DENIED_PURGES
.lock()
.unwrap_or_else(|poisoned| poisoned.into_inner());
if denied.len() >= OBJECT_LOCK_DENIED_PURGE_CACHE_MAX {
denied.retain(|_, entry| entry.denied_at.elapsed() < OBJECT_LOCK_DENIED_PURGE_BACKOFF);
}
let key = object_lock_denied_purge_key(dobj);
if denied.len() < OBJECT_LOCK_DENIED_PURGE_CACHE_MAX || denied.contains_key(&key) {
let entry = denied.entry(key).or_insert_with(|| ObjectLockDeniedPurge {
denied_at: std::time::Instant::now(),
denied_arns: HashSet::new(),
});
entry.denied_at = std::time::Instant::now();
entry.denied_arns.insert(arn.to_string());
}
// Still full after dropping expired entries: skip recording — the purge
// then simply keeps retrying, which is the pre-#6850 behavior.
}
/// Whether a heal requeue of this delete can only reach targets that denied
/// it under object lock within the backoff window. A target the entry does
/// not cover (another peer, or one whose denial expired) keeps the requeue
/// flowing — suppressing it would delay a purge that could succeed there.
pub(crate) fn object_lock_denied_purge_backoff_active(dobj: &DeletedObjectReplicationInfo) -> bool {
let key = object_lock_denied_purge_key(dobj);
let mut denied = OBJECT_LOCK_DENIED_PURGES
.lock()
.unwrap_or_else(|poisoned| poisoned.into_inner());
match denied.get(&key) {
Some(entry) if entry.denied_at.elapsed() < OBJECT_LOCK_DENIED_PURGE_BACKOFF => {
let admitted = dobj.admitted_target_arns();
!admitted.is_empty() && admitted.iter().all(|arn| entry.denied_arns.contains(arn))
}
Some(_) => {
denied.remove(&key);
false
}
None => false,
}
}
const REPLICA_ETAG_VERIFY_ENV: &str = "RUSTFS_REPLICATION_REPLICA_ETAG_VERIFY";
/// Escape hatch for a target whose 32-hex ETags are legitimately not the
/// content MD5 (e.g. a gateway hashing its own ciphertext without announcing
/// SSE in the response) — such a target would otherwise fail every object.
fn replica_etag_verification_enabled() -> bool {
std::env::var(REPLICA_ETAG_VERIFY_ENV)
.map(|v| !(v.eq_ignore_ascii_case("false") || v == "0"))
.unwrap_or(true)
}
/// A 200 from the target is not proof the replica holds the source bytes: a
/// target that stores a transformed payload (e.g. undecoded `aws-chunked`
/// frames, #6853) returns the ETag of what it actually wrote. Reporting
/// COMPLETED over such a replica is silent corruption, so a decidable
/// mismatch fails the replication instead. An SSE-C ciphertext passthrough
/// transfer is exempt: the wire bytes are ciphertext while the source ETag is
/// the plaintext MD5, and that path has its own HEAD-back audit.
fn verify_single_part_replica(
object_info: &ObjectInfo,
response: &RemotePutObjectResponse,
ciphertext_passthrough: bool,
) -> std::result::Result<(), std::io::Error> {
if ciphertext_passthrough || !replica_etag_verification_enabled() {
return Ok(());
}
if single_part_replica_etag_mismatch(object_info.etag.as_deref(), response.etag.as_deref()) {
// The differing ETags go into the structured log; the error message
// stays constant so same-cause failures bucket together downstream.
warn!(
event = EVENT_RESYNC_TARGET_OPERATION_FAILED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
bucket = %object_info.bucket,
object = %object_info.name,
source_etag = ?object_info.etag,
replica_etag = ?response.etag,
operation = "verify_replica_etag",
"Replication target operation failed"
);
return Err(std::io::Error::other(REPLICA_ETAG_MISMATCH_ERROR));
}
Ok(())
}
const REPLICA_ETAG_MISMATCH_ERROR: &str = "replica etag mismatch: the target persisted different bytes than were sent";
fn audit_target_version_identity(tgt_client: &TargetClient, source_version_id: &str, assigned_version_id: Option<&str>) {
if !version_identity_drifted(source_version_id, assigned_version_id) {
return;
@@ -2708,19 +2827,42 @@ async fn replicate_delete_to_target(dobj: &DeletedObjectReplicationInfo, tgt_cli
}
}
Err(e) => {
warn!(
event = EVENT_RESYNC_TARGET_OPERATION_FAILED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
bucket = tgt_client.bucket,
object = dobj.delete_object.object_name,
version_id = ?version_id,
delete_marker = dobj.delete_object.delete_marker,
is_version_purge,
error = %e,
operation = "replicate_delete_to_target",
"Replication target operation failed"
);
let object_lock_denied = is_version_purge && is_object_lock_denied_delete(e.code.as_deref(), e.message.as_deref());
if object_lock_denied {
// Terminal for as long as the lock holds: the peer retains
// this version and replication carries no governance bypass
// (#6850), so the sites stay diverged until the retention or
// legal hold on the replica lapses. Surface it loudly instead
// of letting a silent failed counter and a hot heal-retry
// loop stand in for the divergence.
record_object_lock_denied_purge(dobj, &tgt_client.arn);
error!(
event = EVENT_REPLICATION_PURGE_OBJECT_LOCK_DENIED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
bucket = tgt_client.bucket,
object = dobj.delete_object.object_name,
version_id = ?version_id,
arn = %tgt_client.arn,
error = %e,
operation = "replicate_delete_to_target",
"Replicated version purge denied by object lock on the target; the sites stay diverged until the lock lapses"
);
} else {
warn!(
event = EVENT_RESYNC_TARGET_OPERATION_FAILED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
bucket = tgt_client.bucket,
object = dobj.delete_object.object_name,
version_id = ?version_id,
delete_marker = dobj.delete_object.delete_marker,
is_version_purge,
error = %e,
operation = "replicate_delete_to_target",
"Replication target operation failed"
);
}
rinfo.error = Some(e.to_string());
if !is_version_purge {
rinfo.replication_status = ReplicationStatusType::Failed;
@@ -3274,14 +3416,15 @@ impl ReplicateObjectInfoExt for ReplicateObjectInfo {
let result = tgt_client
.put_object(&tgt_client.bucket, &object, transfer_size, byte_stream, &put_opts)
.await
.map(|assigned_version_id| {
.map_err(|e| std::io::Error::other(e.to_string()))
.and_then(|response| {
audit_target_version_identity(
&tgt_client,
&put_opts.internal.source_version_id,
assigned_version_id.as_deref(),
)
})
.map_err(|e| std::io::Error::other(e.to_string()));
response.version_id.as_deref(),
);
verify_single_part_replica(&object_info, &response, obj_opts.raw_data_movement_read)
});
result.err()
} {
rinfo.replication_status = ReplicationStatusType::Failed;
@@ -3942,14 +4085,15 @@ async fn replicate_all_payload_to_target<S: ReplicationObjectIO>(
.tgt_client
.put_object(&ctx.tgt_client.bucket, ctx.object, ctx.transfer_size, byte_stream, &ctx.put_opts)
.await
.map(|assigned_version_id| {
.map_err(|e| std::io::Error::other(e.to_string()))
.and_then(|response| {
audit_target_version_identity(
ctx.tgt_client,
&ctx.put_opts.internal.source_version_id,
assigned_version_id.as_deref(),
)
})
.map_err(|e| std::io::Error::other(e.to_string()));
response.version_id.as_deref(),
);
verify_single_part_replica(ctx.object_info, &response, ctx.obj_opts.raw_data_movement_read)
});
result.err()
}
}
@@ -4036,28 +4180,132 @@ async fn replicate_object_with_multipart<S: ReplicationObjectIO>(ctx: MultipartR
let arn = ctx.arn;
let result = replicate_multipart_parts_and_complete(ctx, &upload_id).await;
abort_multipart_on_failure(result, dst_bucket, object, &upload_id, arn, || async {
cli.abort_multipart_upload(dst_bucket, object, &upload_id).await
})
abort_multipart_on_failure(
result,
dst_bucket,
object,
&upload_id,
arn,
|| async { cli.abort_multipart_upload(dst_bucket, object, &upload_id).await },
|| {
schedule_replication_abort_retry(
cli.clone(),
dst_bucket.to_string(),
object.to_string(),
upload_id.clone(),
arn.to_string(),
)
},
)
.await
}
const REPLICATION_ABORT_RETRY_ATTEMPTS: u32 = 5;
const REPLICATION_ABORT_RETRY_INITIAL_DELAY_SECS: u64 = 30;
/// The immediate abort usually fails for the same reason the transfer did —
/// the target is unreachable — and MRF only retries the *object*: every replay
/// mints a fresh upload id, so a failed abort would leak its upload on the
/// target forever (#6854). Retry the abort on a detached, bounded backoff
/// (~30s..8m) so it lands once the target comes back; an upload the target no
/// longer knows counts as cleaned up.
fn schedule_replication_abort_retry(cli: Arc<TargetClient>, dst_bucket: String, object: String, upload_id: String, arn: String) {
tokio::spawn(async move {
let mut delay_secs = REPLICATION_ABORT_RETRY_INITIAL_DELAY_SECS;
for attempt in 1..=REPLICATION_ABORT_RETRY_ATTEMPTS {
tokio::time::sleep(tokio::time::Duration::from_secs(delay_secs)).await;
delay_secs = delay_secs.saturating_mul(2);
match cli.abort_multipart_upload(&dst_bucket, &object, &upload_id).await {
Ok(()) => {
info!(
event = EVENT_REPLICATION_ABORT_RETRY_RESOLVED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
target_bucket = %dst_bucket,
object = %object,
arn = %arn,
upload_id = %upload_id,
operation = "abort_multipart_upload_retry",
attempt,
"Replication abort retry cleaned up the orphaned upload"
);
return;
}
Err(err) if target_upload_already_removed(&err) => {
info!(
event = EVENT_REPLICATION_ABORT_RETRY_RESOLVED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
target_bucket = %dst_bucket,
object = %object,
arn = %arn,
upload_id = %upload_id,
operation = "abort_multipart_upload_retry",
attempt,
"Replication abort retry found the upload already removed"
);
return;
}
Err(err) => {
warn!(
event = EVENT_RESYNC_TARGET_OPERATION_FAILED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
target_bucket = %dst_bucket,
object = %object,
arn = %arn,
upload_id = %upload_id,
operation = "abort_multipart_upload_retry",
attempt,
error = %err,
"Replication target operation failed"
);
}
}
}
// Terminal: the upload id stays in the log so an operator can reap it
// with list-multipart-uploads/abort by hand (the #6840 contract).
warn!(
event = EVENT_RESYNC_TARGET_OPERATION_FAILED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
target_bucket = %dst_bucket,
object = %object,
arn = %arn,
upload_id = %upload_id,
operation = "abort_multipart_upload_retry",
result = "gave_up",
"Replication abort retries exhausted; the incomplete upload remains on the target"
);
});
}
/// AWS answers an abort for an unknown upload with `NoSuchUpload`; that means
/// the orphan is gone (aborted elsewhere or expired), which is the goal state.
fn target_upload_already_removed(err: &S3ClientError) -> bool {
err.code.as_deref() == Some("NoSuchUpload")
}
/// Best-effort abort of the target-side multipart upload once the transfer has
/// failed past CreateMultipartUpload; without it every failed attempt leaves an
/// invisible incomplete upload on the target that keeps billing for its parts.
/// The abort outcome never replaces the transfer error: an abort failure is
/// only logged and `result` is returned as-is.
async fn abort_multipart_on_failure<F, Fut>(
async fn abort_multipart_on_failure<F, Fut, R>(
result: std::io::Result<()>,
dst_bucket: &str,
object: &str,
upload_id: &str,
arn: &str,
abort: F,
schedule_abort_retry: R,
) -> std::io::Result<()>
where
F: FnOnce() -> Fut,
Fut: std::future::Future<Output = std::result::Result<(), S3ClientError>>,
R: FnOnce(),
{
if result.is_ok() {
return result;
@@ -4075,6 +4323,9 @@ where
error = %abort_err,
"Replication target operation failed"
);
if !target_upload_already_removed(&abort_err) {
schedule_abort_retry();
}
}
result
}
@@ -5354,27 +5605,72 @@ mod tests {
assert!(!resync_state_accepts_update(&current, &stale));
}
#[test]
fn object_lock_denied_purge_backoff_tracks_version_and_target() {
let denied = DeletedObjectReplicationInfo {
bucket: "worm-backoff-test-bucket".to_string(),
target_arn: "arn:rustfs:replication::worm-test:t1".to_string(),
delete_object: ReplicationDeletedObject {
object_name: "locked-object".to_string(),
version_id: Some(uuid::Uuid::new_v4()),
..Default::default()
},
..Default::default()
};
assert!(!object_lock_denied_purge_backoff_active(&denied));
record_object_lock_denied_purge(&denied, "arn:rustfs:replication::worm-test:t1");
assert!(object_lock_denied_purge_backoff_active(&denied));
// A requeue that can also reach a target this denial does not cover
// must keep flowing: the purge may succeed there.
let mut other_target = denied.clone();
other_target.target_arn = "arn:rustfs:replication::worm-test:t2".to_string();
assert!(!object_lock_denied_purge_backoff_active(&other_target));
// A different version of the same object must not be suppressed.
let mut other_version = denied;
other_version.delete_object.version_id = Some(uuid::Uuid::new_v4());
assert!(!object_lock_denied_purge_backoff_active(&other_version));
}
#[tokio::test]
async fn abort_multipart_on_failure_skips_abort_when_transfer_succeeded() {
let aborted = Arc::new(AtomicBool::new(false));
let flag = aborted.clone();
let retry_scheduled = Arc::new(AtomicBool::new(false));
let retry_flag = retry_scheduled.clone();
let result = abort_multipart_on_failure(Ok(()), "dst-bucket", "obj", "upload-1", "arn:dest", move || async move {
flag.store(true, Ordering::SeqCst);
Ok(())
})
let result = abort_multipart_on_failure(
Ok(()),
"dst-bucket",
"obj",
"upload-1",
"arn:dest",
move || async move {
flag.store(true, Ordering::SeqCst);
Ok(())
},
move || retry_flag.store(true, Ordering::SeqCst),
)
.await;
assert!(result.is_ok());
assert!(!aborted.load(Ordering::SeqCst));
assert!(!retry_scheduled.load(Ordering::SeqCst));
}
#[tokio::test]
async fn abort_multipart_on_failure_aborts_and_keeps_transfer_error() {
let aborted = Arc::new(AtomicBool::new(false));
let flag = aborted.clone();
let retry_scheduled = Arc::new(AtomicBool::new(false));
let retry_flag = retry_scheduled.clone();
// The abort itself failing must not mask the transfer error.
// The abort itself failing must not mask the transfer error, and a
// failed abort must hand the upload id to the retry schedule (#6854):
// the object itself is re-replicated under a fresh upload id, so
// nothing else will ever abort this one.
let result = abort_multipart_on_failure(
Err(std::io::Error::other("transfer failed")),
"dst-bucket",
@@ -5385,10 +5681,34 @@ mod tests {
flag.store(true, Ordering::SeqCst);
Err(S3ClientError::new("abort failed"))
},
move || retry_flag.store(true, Ordering::SeqCst),
)
.await;
assert!(aborted.load(Ordering::SeqCst));
assert!(retry_scheduled.load(Ordering::SeqCst));
assert_eq!(result.unwrap_err().to_string(), "transfer failed");
}
#[tokio::test]
async fn abort_multipart_on_failure_does_not_retry_a_gone_upload() {
let retry_scheduled = Arc::new(AtomicBool::new(false));
let retry_flag = retry_scheduled.clone();
let result = abort_multipart_on_failure(
Err(std::io::Error::other("transfer failed")),
"dst-bucket",
"obj",
"upload-1",
"arn:dest",
|| async { Err(S3ClientError::with_metadata("gone", None, Some("NoSuchUpload".to_string()), None)) },
move || retry_flag.store(true, Ordering::SeqCst),
)
.await;
// NoSuchUpload means the orphan no longer exists; retrying would only
// produce noise.
assert!(!retry_scheduled.load(Ordering::SeqCst));
assert_eq!(result.unwrap_err().to_string(), "transfer failed");
}
}
@@ -36,8 +36,8 @@ use time::OffsetDateTime;
use time::format_description::well_known::Rfc3339;
pub(crate) use crate::bucket::bucket_target_sys::{
AdvancedPutOptions, HeadObjectSdkError, PutObjectOptions, PutObjectPartOptions, RemoveObjectOptions, S3ClientError,
TargetClient, resolve_read_api_version_id,
AdvancedPutOptions, HeadObjectSdkError, PutObjectOptions, PutObjectPartOptions, RemotePutObjectResponse, RemoveObjectOptions,
S3ClientError, TargetClient, resolve_read_api_version_id,
};
#[cfg(test)]
pub(crate) use crate::bucket::target::BucketTarget;
@@ -1307,6 +1307,32 @@ pub fn verify_tonic_mutation_body_digest<T>(request: &tonic::Request<T>, canonic
verify_tonic_mutation_body_digest_with_strictness(request, canonical_body, internode_rpc_body_digest_strict())
}
/// Verify a non-disk mutation without accepting a newly-generated unsigned v2 body.
///
/// The disk mutation lane has a rolling-upgrade exception for `UNSIGNED-PAYLOAD`
/// while peer replay-cache capability is being discovered. Historical v2 peers
/// used the fixed `unsigned` nonce before body-digest rollout; preserve that
/// exact marker for mixed-version compatibility, but reject unsigned v2
/// requests that omit it or present a different nonce.
pub fn verify_tonic_mutation_body_digest_reject_unsigned<T>(
request: &tonic::Request<T>,
canonical_body: &[u8],
) -> std::io::Result<()> {
let version = request
.metadata()
.get(RPC_AUTH_VERSION_HEADER)
.and_then(|value| value.to_str().ok());
let digest = request
.metadata()
.get(RPC_CONTENT_SHA256_HEADER)
.and_then(|value| value.to_str().ok());
let nonce = request.metadata().get(RPC_NONCE_HEADER).and_then(|value| value.to_str().ok());
if version == Some(RPC_AUTH_VERSION_V2) && digest == Some(UNSIGNED_PAYLOAD) && nonce != Some("unsigned") {
return Err(std::io::Error::other("RPC mutation requires a body-bound v2 signature"));
}
verify_tonic_mutation_body_digest(request, canonical_body)
}
/// [`verify_tonic_mutation_body_digest`] with the strict gate injected as a parameter, so both
/// rollout postures are unit-testable without racing on process-global environment variables.
fn verify_tonic_mutation_body_digest_with_strictness<T>(
+2 -2
View File
@@ -39,8 +39,8 @@ pub use http_auth::{
sign_tonic_rpc_response_proof, tonic_boot_epoch_challenge, tonic_boot_epoch_response_headers, tonic_rpc_auth_failure_reason,
verify_ns_scanner_capability, verify_ns_scanner_capability_with_tier_registry_generation, verify_put_file_auth_trailer,
verify_put_file_capability, verify_rpc_signature, verify_tonic_boot_epoch_response, verify_tonic_canonical_body_digest,
verify_tonic_mutation_body_digest, verify_tonic_rpc_response_proof, verify_tonic_rpc_signature,
verify_tonic_rpc_signature_with_bootstrap,
verify_tonic_mutation_body_digest, verify_tonic_mutation_body_digest_reject_unsigned, verify_tonic_rpc_response_proof,
verify_tonic_rpc_signature, verify_tonic_rpc_signature_with_bootstrap,
};
#[cfg(test)]
pub(crate) use internode_data_transport::TcpHttpInternodeDataTransport;
+163 -17
View File
@@ -109,7 +109,10 @@ static USAGE_MEMORY_GENERATION: AtomicU64 = AtomicU64::new(0);
/// strictly tighter than beta.11 (usage treated as 0) and strictly more
/// available than a blanket 503. The fallback applies to any window without
/// authoritative usage, not only pre-v2 upgrades; the values always come from
/// the last persisted scanner output. Loads go through the TTL-bounded
/// the last persisted scanner output — pre-discard sizes of the
/// authoritative snapshot first, backfilled per bucket from the observed
/// (nonconverged) snapshot for buckets no authoritative cycle has covered
/// yet (issue #6852). Loads go through the TTL-bounded
/// snapshot cache, so the quota path adds at most one backend read per
/// [`DATA_USAGE_CACHE_TTL_SECS`] window. Returns `None` for buckets absent
/// from every persisted snapshot — those still fail closed.
@@ -168,7 +171,7 @@ fn fresh_cached_data_usage_snapshot(
fn cache_data_usage_snapshot_result(
cache: &mut Option<CachedDataUsageSnapshot>,
result: Result<(DataUsageInfo, HashMap<String, u64>), Error>,
result: Result<LoadedUsageBaseline, Error>,
loaded_at: tokio::time::Instant,
refresh_generation: u64,
current_generation: u64,
@@ -178,7 +181,19 @@ fn cache_data_usage_snapshot_result(
}
Some(match result {
Ok((info, degraded_baseline)) => {
Ok(LoadedUsageBaseline {
info,
mut degraded_baseline,
observed_unavailable,
}) => {
// A flaky observed read must not shrink quota coverage for a TTL
// window: carry the previous refresh's baseline entries forward,
// letting the fresh (authoritative) values win where they exist.
if observed_unavailable && let Some(previous) = cache.as_ref() {
for (bucket, size) in &previous.degraded_baseline {
degraded_baseline.entry(bucket.clone()).or_insert(*size);
}
}
*cache = Some(CachedDataUsageSnapshot {
info: Some(info.clone()),
loaded_at,
@@ -1113,24 +1128,78 @@ async fn load_data_usage_snapshot(store: Arc<ECStore>) -> Result<(DataUsageInfo,
/// Load data usage info from backend storage
#[instrument(skip(store))]
pub async fn load_data_usage_from_backend(store: Arc<ECStore>) -> Result<DataUsageInfo, Error> {
Ok(load_data_usage_from_backend_with_baseline(store).await?.0)
Ok(load_data_usage_from_backend_with_baseline(store).await?.info)
}
/// One refresh of the persisted usage snapshot plus the quota-admission
/// baseline derived from it.
struct LoadedUsageBaseline {
info: DataUsageInfo,
degraded_baseline: HashMap<String, u64>,
/// True when the observed snapshot could not be read (a transport error,
/// not absence): the cached loader then carries the previous refresh's
/// baseline entries forward instead of shrinking quota coverage for a
/// whole TTL window over one flaky read.
observed_unavailable: bool,
}
/// Like [`load_data_usage_from_backend`], but also returns the pre-discard
/// per-bucket sizes so the cached loader can retain them as the degraded
/// quota-admission baseline (issue #5716).
async fn load_data_usage_from_backend_with_baseline(store: Arc<ECStore>) -> Result<(DataUsageInfo, HashMap<String, u64>), Error> {
let (data_usage_info, source) = load_data_usage_snapshot(store).await?;
Ok(normalize_loaded_data_usage(data_usage_info, source.is_authoritative()).await)
async fn load_data_usage_from_backend_with_baseline(store: Arc<ECStore>) -> Result<LoadedUsageBaseline, Error> {
let (loaded_snapshot, source) = load_data_usage_snapshot(store.clone()).await?;
// The observed-newness gate below compares against the snapshot as
// persisted, before normalization demotes or discards anything.
let authoritative_as_persisted = loaded_snapshot.clone();
let (info, mut degraded_baseline) = normalize_loaded_data_usage(loaded_snapshot, source.is_authoritative()).await;
// A bucket without a converged scanner cycle behind it — a freshly joined
// replica whose every cycle is superseded by the sustained replication
// write stream, or a bucket created after the last converged cycle on a
// busy site (#6852) — has no authoritative size, and quota admission
// fails its writes closed indefinitely. The observed (nonconverged)
// snapshot those superseded cycles still publish is the only grounded
// usage in that window, so it backfills buckets the loaded baseline does
// not cover; a value already in the baseline always wins. The newness
// gate ties the observation to this exact authoritative snapshot, so a
// stale observed object left behind by an earlier incarnation (e.g. a
// deleted and recreated bucket) cannot inject ghost usage. Loads sit
// behind the same TTL cache as the snapshot itself, so this adds at most
// one backend read per TTL window.
let mut observed_unavailable = false;
match load_observed_data_usage_snapshot(store).await {
Ok(Some(observed)) if observed_data_usage_is_newer(&observed, &authoritative_as_persisted) => {
backfill_degraded_baseline_from_observed(&mut degraded_baseline, &observed);
}
Ok(_) => {}
Err(_) => observed_unavailable = true,
}
Ok(LoadedUsageBaseline {
info,
degraded_baseline,
observed_unavailable,
})
}
async fn load_observed_data_usage_snapshot(store: Arc<ECStore>) -> Option<DataUsageInfo> {
/// Fill quota-baseline gaps from an observed (nonconverged) snapshot without
/// overriding any bucket the authoritative baseline already covers.
fn backfill_degraded_baseline_from_observed(degraded_baseline: &mut HashMap<String, u64>, observed: &DataUsageInfo) {
for (bucket, usage) in &observed.buckets_usage {
degraded_baseline.entry(bucket.clone()).or_insert(usage.size);
}
}
/// `Ok(None)` means the observed snapshot is absent or invalid (a settled
/// answer); `Err` means it could not be read at all, so the caller may keep
/// using what it learned from a previous read.
async fn load_observed_data_usage_snapshot(store: Arc<ECStore>) -> Result<Option<DataUsageInfo>, Error> {
let data = match read_config_preserve_empty(store, &DATA_USAGE_OBSERVED_OBJ_NAME_PATH).await {
Ok(data) => data,
Err(Error::ConfigNotFound) => return None,
Err(Error::ConfigNotFound) => return Ok(None),
Err(err) => {
record_usage_snapshot_failure("read_observed", DATA_USAGE_OBSERVED_OBJ_NAME_PATH.as_str(), &err);
return None;
return Err(err);
}
};
@@ -1139,7 +1208,7 @@ async fn load_observed_data_usage_snapshot(store: Arc<ECStore>) -> Option<DataUs
if info.usage_snapshot_converged == Some(false)
&& (info.is_complete_bucket_usage_snapshot() || info.is_valid_partial_snapshot()) =>
{
Some(info)
Ok(Some(info))
}
Ok(_) => {
error!(
@@ -1150,11 +1219,11 @@ async fn load_observed_data_usage_snapshot(store: Arc<ECStore>) -> Option<DataUs
object = %DATA_USAGE_OBSERVED_OBJ_NAME_PATH.as_str(),
"observed data usage snapshot was not a structurally complete nonconverged view"
);
None
Ok(None)
}
Err(err) => {
record_usage_snapshot_decode_failure("parse_observed", DATA_USAGE_OBSERVED_OBJ_NAME_PATH.as_str(), &err);
None
Ok(None)
}
}
}
@@ -1212,7 +1281,9 @@ fn merge_partial_observation_for_admin(mut authoritative: DataUsageInfo, observe
async fn load_admin_data_usage_from_backend(store: Arc<ECStore>) -> Result<DataUsageInfo, Error> {
let (authoritative, source) = load_data_usage_snapshot(store.clone()).await?;
let observed = load_observed_data_usage_snapshot(store).await;
// For the one-shot admin view a failed observed read degrades to "no
// observation", same as before the read was fallible.
let observed = load_observed_data_usage_snapshot(store).await.ok().flatten();
let (selected, selected_is_current_format) =
select_admin_data_usage_snapshot(authoritative, source.is_authoritative(), observed);
Ok(normalize_loaded_data_usage(selected, selected_is_current_format).await.0)
@@ -1375,7 +1446,11 @@ pub async fn load_admin_data_usage_from_backend_cached(store: Arc<ECStore>) -> R
let refresh_generation = admin_data_usage_snapshot_generation();
let result = load_admin_data_usage_from_backend(store.clone())
.await
.map(|info| (info, HashMap::new()));
.map(|info| LoadedUsageBaseline {
info,
degraded_baseline: HashMap::new(),
observed_unavailable: false,
});
let loaded_at = tokio::time::Instant::now();
let mut cache = admin_data_usage_snapshot_cache().write().await;
if let Some(result) = cache_data_usage_snapshot_result(
@@ -2526,6 +2601,37 @@ mod tests {
use std::sync::Arc;
use tokio::{io::AsyncReadExt, sync::Mutex};
#[test]
fn observed_snapshot_only_backfills_baseline_gaps() {
let mut baseline = HashMap::from([("covered".to_string(), 111_u64)]);
let observed = DataUsageInfo {
buckets_usage: HashMap::from([
(
"covered".to_string(),
BucketUsageInfo {
size: 999,
..Default::default()
},
),
(
"replica-only".to_string(),
BucketUsageInfo {
size: 42,
..Default::default()
},
),
]),
..Default::default()
};
backfill_degraded_baseline_from_observed(&mut baseline, &observed);
// The authoritative value must win; only the uncovered bucket (#6852:
// a replica that never landed a converged cycle) is filled in.
assert_eq!(baseline.get("covered"), Some(&111));
assert_eq!(baseline.get("replica-only"), Some(&42));
}
#[derive(Debug, Default)]
struct UsageCasState {
object: Option<(Vec<u8>, u64)>,
@@ -3479,7 +3585,11 @@ mod tests {
let first = cache_data_usage_snapshot_result(
&mut cache,
Ok((expected, HashMap::new())),
Ok(LoadedUsageBaseline {
info: expected,
degraded_baseline: HashMap::new(),
observed_unavailable: false,
}),
loaded_at,
refresh_generation,
data_usage_snapshot_generation(),
@@ -3494,6 +3604,38 @@ mod tests {
assert_snapshot_bucket(&cached, "bucket");
}
#[test]
#[serial]
fn unavailable_observed_read_keeps_previous_baseline_coverage() {
let loaded_at = tokio::time::Instant::now();
let refresh_generation = data_usage_snapshot_generation();
let mut cache = Some(CachedDataUsageSnapshot {
info: Some(data_usage_info_for_test("bucket", 1, 42, SystemTime::UNIX_EPOCH)),
loaded_at,
degraded_baseline: HashMap::from([("observed-only".to_string(), 7_u64), ("covered".to_string(), 1)]),
});
cache_data_usage_snapshot_result(
&mut cache,
Ok(LoadedUsageBaseline {
info: data_usage_info_for_test("bucket", 1, 42, SystemTime::UNIX_EPOCH),
degraded_baseline: HashMap::from([("covered".to_string(), 2_u64)]),
observed_unavailable: true,
}),
loaded_at,
refresh_generation,
data_usage_snapshot_generation(),
)
.expect("an uninterrupted refresh should populate the cache")
.expect("successful load must be returned");
let baseline = &cache.as_ref().expect("cache must be populated").degraded_baseline;
// The bucket only the (now unreadable) observed snapshot covered must
// survive the refresh; the freshly loaded value wins where it exists.
assert_eq!(baseline.get("observed-only"), Some(&7));
assert_eq!(baseline.get("covered"), Some(&2));
}
#[test]
#[serial]
fn cache_invalidation_during_refresh_prevents_stale_snapshot_resurrection() {
@@ -3508,7 +3650,11 @@ mod tests {
let stale_result = cache_data_usage_snapshot_result(
&mut cache,
Ok((data_usage_info_for_test("stale", 1, 42, SystemTime::UNIX_EPOCH), HashMap::new())),
Ok(LoadedUsageBaseline {
info: data_usage_info_for_test("stale", 1, 42, SystemTime::UNIX_EPOCH),
degraded_baseline: HashMap::new(),
observed_unavailable: false,
}),
loaded_at,
refresh_generation,
data_usage_snapshot_generation(),
+75 -45
View File
@@ -574,6 +574,7 @@ pub(crate) struct ParallelReader<R> {
read_timeout: Duration,
verify_reconstruction: bool,
locality_preference_enabled: bool,
demand_bound_lockstep: bool,
// Request-scoped shard buffers keyed by shard index. Keeping ownership in
// `ParallelReader` avoids dropping unused parity/backup slot buffers between stripes.
buffers: ShardBufferPool,
@@ -585,10 +586,8 @@ pub(crate) struct ParallelReader<R> {
// it to the current stripe when it is engaged mid-object (backlog#923).
engaged: SmallVec<[bool; INLINE_SHARD_SLOTS]>,
deferred_handles: Vec<Option<DeferredReaderStripeHandle>>,
// Copy-source hedges use a fresh deferred reader so cancelling a hedge
// never consumes the unopened reader reserved for a later stripe. The
// vector is empty for callers that do not provide a reopen factory (tests
// and the ordinary GET path retain the handle-based behavior).
// Demand-bound hedges use a fresh deferred reader so cancelling a hedge
// never consumes the unopened reader reserved for a later stripe.
deferred_reopeners: Vec<Option<DeferredReaderReopener<R>>>,
stripe_index: usize,
}
@@ -777,9 +776,9 @@ where
// reads all live readers on every stripe — the pre-backlog#923
// behavior. With the gate on, only data slots start engaged; parity is
// engaged on demand, stripe-aligned through its deferred handle.
let data_shards_only = get_lockstep_data_shards_only_enabled();
let demand_bound_lockstep = get_lockstep_data_shards_only_enabled();
let engaged: SmallVec<_> = (0..readers.len())
.map(|index| !data_shards_only || index < e.data_shards)
.map(|index| !demand_bound_lockstep || index < e.data_shards)
.collect();
ParallelReader {
readers,
@@ -793,6 +792,7 @@ where
read_timeout,
verify_reconstruction,
locality_preference_enabled: get_shard_locality_preference_enabled(),
demand_bound_lockstep,
buffers: ShardBufferPool::new(e.data_shards + e.parity_shards),
stripe_state: None,
engaged,
@@ -1275,7 +1275,7 @@ where
/// realigned (no pending deferred handle) is likewise retired instead of
/// being read out of position.
async fn read_lockstep(&mut self, state: &mut StripeReadState) {
if matches!(decode_read_policy(), DecodeReadPolicy::DemandBound) {
if self.demand_bound_lockstep {
self.read_lockstep_demand_bound(state).await;
return;
}
@@ -1531,17 +1531,18 @@ where
}
}
/// Demand-bound lockstep stripe read used by server-side copy sources.
/// Demand-bound data-shards-only lockstep stripe read.
///
/// The ordinary lockstep path can cancel every in-flight reader once it
/// has a quorum because all of its parity readers are already engaged.
/// Copy sources keep parity unopened until a data reader is missing. A
/// hedge therefore has to race the deferred parity reads against the
/// original data reads and may retire the latter only after the parity has
/// produced an actual decode-plus-verification quorum. The futures own
/// their readers so disjoint data/parity slots can be admitted while the
/// other group is still pending; dropping an abandoned future retires its
/// stream without leaving a borrowed slot behind.
/// Copy sources and the data-shards-only rollout gate keep parity unopened
/// until a data reader is missing. A hedge therefore has to race the
/// deferred parity reads against the original data reads and may retire the
/// latter only after parity has produced an actual decode-plus-verification
/// quorum. The futures own their readers so disjoint data/parity slots can
/// be admitted while the other group is still pending; dropping an
/// abandoned future retires its stream without leaving a borrowed slot
/// behind.
async fn read_lockstep_demand_bound(&mut self, state: &mut StripeReadState) {
let num_readers = self.readers.len();
state.reset(num_readers, self.data_shards);
@@ -1576,14 +1577,14 @@ where
let mut completed = 0usize;
let mut failed = 0usize;
let mut first_shard_recorded = false;
let mut active = vec![false; num_readers];
let mut temporary_parity = vec![false; num_readers];
let mut active: ActiveReaders = smallvec![false; num_readers];
let mut temporary_parity: ActiveReaders = smallvec![false; num_readers];
// A deferred parity slot is attempted at most once per stripe. A
// failed disposable hedge keeps its unopened reserve for the next
// stripe, but must not be relaunched in a tight same-stripe retry
// loop (which would defeat the bounded fan-out and amplify a remote
// outage).
let mut attempted_parity = vec![false; num_readers];
let mut attempted_parity: ActiveReaders = smallvec![false; num_readers];
// Once a data reader has returned an error (or was already missing at
// setup), the loss is permanent for lockstep alignment. Use the
// deferred handle and keep parity engaged across subsequent stripes;
@@ -4911,6 +4912,24 @@ mod tests {
/// read timeout even though both parity readers were available to engage.
#[tokio::test]
async fn test_demand_bound_lockstep_hedges_to_deferred_parity_quorum() {
with_decode_read_policy(DecodeReadPolicy::DemandBound, assert_deferred_parity_hedges_slow_data()).await;
}
/// The ordinary GET rollout gate must use the same bounded parity race as
/// CopySource. Leaving it on the legacy lockstep loop deadlocks the hedge:
/// that loop waits for a parity success before cancelling the slow data
/// read, but does not admit deferred parity until after the data read ends.
#[tokio::test]
#[serial_test::serial]
async fn test_data_shards_only_gate_hedges_to_deferred_parity_quorum() {
temp_env::async_with_vars(
[(ENV_RUSTFS_GET_LOCKSTEP_DATA_SHARDS_ONLY_ENABLE, Some("true"))],
assert_deferred_parity_hedges_slow_data(),
)
.await;
}
async fn assert_deferred_parity_hedges_slow_data() {
const NUM_SHARDS: usize = 1;
const BLOCK_SIZE: usize = 64;
const DATA_SHARDS: usize = 2;
@@ -4951,33 +4970,27 @@ mod tests {
];
let erasure = Erasure::new(DATA_SHARDS, PARITY_SHARDS, BLOCK_SIZE);
let (bufs, errs, engaged, readers_remaining) = with_decode_read_policy(DecodeReadPolicy::DemandBound, async {
let mut parallel_reader = ParallelReader::new_with_metrics_path_read_costs_timeout_and_reconstruction_verification(
readers,
erasure,
0,
NUM_SHARDS * BLOCK_SIZE,
None,
vec![ShardReadCost::Unknown; DATA_SHARDS + PARITY_SHARDS],
Duration::from_secs(60),
true,
);
let (bufs, errs) = tokio::time::timeout(Duration::from_secs(2), parallel_reader.read())
.await
.expect("deferred parity must cover a hedged data shard without waiting for read_timeout");
(
bufs,
errs,
parallel_reader.engaged.clone(),
parallel_reader.readers.iter().map(Option::is_some).collect::<Vec<_>>(),
)
})
.await;
let mut parallel_reader = ParallelReader::new_with_metrics_path_read_costs_timeout_and_reconstruction_verification(
readers,
erasure,
0,
NUM_SHARDS * BLOCK_SIZE,
None,
vec![ShardReadCost::Unknown; DATA_SHARDS + PARITY_SHARDS],
Duration::from_secs(60),
true,
);
let (bufs, errs) = tokio::time::timeout(Duration::from_secs(2), parallel_reader.read())
.await
.expect("deferred parity must cover a hedged data shard without waiting for read_timeout");
assert!(matches!(&errs[0], Some(DiskError::Io(err)) if err.kind() == ErrorKind::TimedOut));
assert_eq!(bufs.iter().filter(|buf| buf.is_some()).count(), DATA_SHARDS + 1);
assert_eq!(engaged.as_slice(), &[true, true, true, true]);
assert_eq!(readers_remaining, vec![false, true, true, true]);
assert_eq!(parallel_reader.engaged.as_slice(), &[true, true, true, true]);
assert_eq!(
parallel_reader.readers.iter().map(Option::is_some).collect::<Vec<_>>(),
vec![false, true, true, true]
);
}
/// A fast data failure must admit deferred parity immediately. There is
@@ -5046,6 +5059,24 @@ mod tests {
#[tokio::test]
async fn test_demand_bound_canceled_hedge_preserves_deferred_parity_for_next_stripe() {
with_decode_read_policy(
DecodeReadPolicy::DemandBound,
assert_canceled_hedge_preserves_deferred_parity_for_next_stripe(),
)
.await;
}
#[tokio::test]
#[serial_test::serial]
async fn test_data_shards_only_gate_canceled_hedge_preserves_deferred_parity_for_next_stripe() {
temp_env::async_with_vars(
[(ENV_RUSTFS_GET_LOCKSTEP_DATA_SHARDS_ONLY_ENABLE, Some("true"))],
assert_canceled_hedge_preserves_deferred_parity_for_next_stripe(),
)
.await;
}
async fn assert_canceled_hedge_preserves_deferred_parity_for_next_stripe() {
const BLOCK_SIZE: usize = 64;
const DATA_SHARDS: usize = 2;
const PARITY_SHARDS: usize = 2;
@@ -5094,7 +5125,7 @@ mod tests {
Some(BitrotReader::new(TestShardReader::Pending, SHARD_SIZE, hash_algo, false)),
];
let (first_parity_reserved, second_result) = with_decode_read_policy(DecodeReadPolicy::DemandBound, async {
let (first_parity_reserved, second_result) = {
let erasure = Erasure::new(DATA_SHARDS, PARITY_SHARDS, BLOCK_SIZE);
let mut parallel_reader = ParallelReader::new_with_metrics_path_read_timeout_and_reconstruction_verification(
readers,
@@ -5155,8 +5186,7 @@ mod tests {
parallel_reader.readers[2].is_some() && parallel_reader.readers[3].is_some(),
(third_buffers, third_errors),
)
})
.await;
};
assert!(first_parity_reserved);
assert_eq!(parity_calls.load(Ordering::SeqCst), PARITY_SHARDS * 2);
+308 -21
View File
@@ -37,6 +37,8 @@ use super::super::ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS;
#[cfg(test)]
use super::super::ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX;
#[cfg(test)]
use super::super::ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE;
#[cfg(test)]
use super::super::get_metadata_slowtail_fault_delay;
use super::super::{
Bytes, CHECK_PART_DISK_NOT_FOUND, DeleteOptions, DiskError, DiskStore, EVENT_SET_DISK_RENAME_TAIL_DRAIN_FAILED,
@@ -49,10 +51,11 @@ use super::super::{
capacity_scope_from_disks, coding, collect_inline_data_shard_fileinfos_by_index_or_reason, current_dirty_generation, debug,
disk, file_info_is_valid_for_metadata, get_metadata_slowtail_fault_request, info, inline_erasure_shard_file_offset,
inline_erasure_shard_size, is_err_object_not_found, is_err_version_not_found, is_get_metadata_data_read_early_stop_enabled,
is_get_metadata_early_stop_bounded_fanout_enabled, is_get_metadata_early_stop_enabled, is_object_dangling,
is_version_early_stop_enabled, issue3031_diag_enabled, join_all, join_errs, log_multipart_write_quorum_failure,
merge_file_meta_versions, path_join_buf, record_global_dirty_scope, reduce_read_quorum_errs, reduce_write_quorum_errs,
send_heal_request_with_admission, should_prevent_write, to_object_err, try_read_inline_data_shards_direct, warn,
is_get_metadata_early_stop_bounded_fanout_enabled, is_get_metadata_early_stop_enabled,
is_get_metadata_non_inline_data_read_early_stop_enabled, is_object_dangling, is_version_early_stop_enabled,
issue3031_diag_enabled, join_all, join_errs, log_multipart_write_quorum_failure, merge_file_meta_versions, path_join_buf,
record_global_dirty_scope, reduce_read_quorum_errs, reduce_write_quorum_errs, send_heal_request_with_admission,
should_prevent_write, to_object_err, try_read_inline_data_shards_direct, warn,
};
#[cfg(test)]
use crate::bucket::lifecycle::lifecycle::TRANSITION_COMPLETE;
@@ -687,6 +690,10 @@ pub(in crate::set_disk) struct MetadataQuorumAccumulator {
pub(in crate::set_disk) hard_errors: usize,
pub(in crate::set_disk) candidate: Option<FileInfo>,
pub(in crate::set_disk) candidate_votes: usize,
// Bitset of shard indexes whose metadata matches the candidate. Erasure
// layouts are capped at 16 shards, so this stays allocation-free on the
// GET metadata hot path.
candidate_shard_mask: u16,
pub(in crate::set_disk) conflicting_metadata: bool,
pub(in crate::set_disk) delete_marker_seen: bool,
pub(in crate::set_disk) delete_marker_candidates: Vec<(FileInfo, usize)>,
@@ -708,6 +715,7 @@ impl MetadataQuorumAccumulator {
hard_errors: 0,
candidate: None,
candidate_votes: 0,
candidate_shard_mask: 0,
conflicting_metadata: false,
delete_marker_seen: false,
delete_marker_candidates: Vec::new(),
@@ -723,6 +731,14 @@ impl MetadataQuorumAccumulator {
}
pub(in crate::set_disk) fn observe_file_info(&mut self, file_info: &FileInfo) {
self.observe_file_info_with_index(None, file_info);
}
pub(in crate::set_disk) fn observe_file_info_at(&mut self, disk_index: usize, file_info: &FileInfo) {
self.observe_file_info_with_index(Some(disk_index), file_info);
}
fn observe_file_info_with_index(&mut self, disk_index: Option<usize>, file_info: &FileInfo) {
if !file_info_is_valid_for_metadata(file_info) {
self.hard_errors = self.hard_errors.saturating_add(1);
return;
@@ -762,6 +778,11 @@ impl MetadataQuorumAccumulator {
match &self.candidate {
Some(candidate) if metadata_early_stop_candidate_matches(candidate, file_info) => {
self.candidate_votes = self.candidate_votes.saturating_add(1);
if let Some(disk_index) = disk_index
&& let Some(bit) = Self::candidate_shard_bit(candidate, file_info, disk_index)
{
self.candidate_shard_mask |= bit;
}
}
Some(_) => {
self.conflicting_metadata = true;
@@ -769,10 +790,38 @@ impl MetadataQuorumAccumulator {
None => {
self.candidate = Some(file_info.clone());
self.candidate_votes = 1;
if let Some(disk_index) = disk_index
&& let Some(bit) = Self::candidate_shard_bit(file_info, file_info, disk_index)
{
self.candidate_shard_mask |= bit;
}
}
}
}
fn candidate_shard_bit(candidate: &FileInfo, file_info: &FileInfo, disk_index: usize) -> Option<u16> {
let &erasure_index = candidate.erasure.distribution.get(disk_index)?;
if erasure_index == 0 || erasure_index > u16::BITS as usize || file_info.erasure.index != erasure_index {
return None;
}
Some(1u16 << (erasure_index - 1))
}
pub(in crate::set_disk) fn candidate_has_read_reserve(&self) -> bool {
self.candidate_read_reserve_target()
.is_some_and(|required| self.candidate_shard_mask.count_ones() as usize >= required)
}
pub(in crate::set_disk) fn candidate_read_reserve_target(&self) -> Option<usize> {
let candidate = self.candidate.as_ref()?;
Some(
candidate
.erasure
.data_blocks
.saturating_add(usize::from(candidate.erasure.parity_blocks > 0)),
)
}
pub(in crate::set_disk) fn observe_error(&mut self, err: &DiskError) {
match err {
DiskError::FileNotFound | DiskError::VolumeNotFound => {
@@ -1083,6 +1132,23 @@ fn data_read_early_stop_inline_candidate_miss_reason(candidate: &FileInfo) -> Op
None
}
fn non_inline_data_read_candidate_is_safe(candidate: &FileInfo) -> bool {
if candidate.inline_data()
|| candidate.is_compressed()
|| candidate.is_remote()
|| candidate
.metadata
.keys()
.any(|key| rustfs_utils::http::is_object_encryption_marker(key))
|| candidate.parts.len() != 1
{
return false;
}
candidate.has_valid_erasure_geometry()
}
const NON_INLINE_SINGLE_PENDING_HEDGE_DELAY: Duration = Duration::from_millis(100);
fn data_read_inline_missing_shards_are_pending(
candidate: &FileInfo,
parts_metadata: &[FileInfo],
@@ -1929,14 +1995,10 @@ pub(in crate::set_disk) fn fill_deferred_bitrot_readers(
return;
}
// Only CopySource uses disposable, stripe-aligned reopeners. Ordinary GET
// readers use the existing deferred handle and should not retain one
// heap-allocated closure (plus cloned path/disk state) for every parity
// slot.
let copy_source_demand_bound = matches!(
crate::set_disk::get_object_read_policy(),
crate::set_disk::GetObjectReadPolicy::CopySource
);
// Every demand-bound lockstep reader needs a disposable, stripe-aligned
// reopener. Otherwise a recovered slow data read can cancel and consume
// the only parity reserve needed by a later degraded stripe.
let demand_bound_lockstep = crate::erasure::coding::decode::get_lockstep_data_shards_only_enabled();
for idx in 0..disks.len() {
if setup.attempted[idx] {
@@ -1951,7 +2013,7 @@ pub(in crate::set_disk) fn fill_deferred_bitrot_readers(
let disk = disks[idx].clone();
let data_dir = files[idx].data_dir.unwrap_or_default();
let path = format!("{object}/{data_dir}/part.{part_number}");
let reopener = copy_source_demand_bound.then(|| {
let reopener = demand_bound_lockstep.then(|| {
deferred_reader_reopener(
inline_data.clone(),
disk.clone(),
@@ -1992,7 +2054,7 @@ pub(in crate::set_disk) fn fill_deferred_bitrot_readers(
// ready/error bookkeeping that quorum decisions rely on is left untouched.
// Gate off (default): keep the eagerly opened parity readers exactly as
// before — the lockstep path reads them on every stripe.
if !crate::erasure::coding::decode::get_lockstep_data_shards_only_enabled() {
if !demand_bound_lockstep {
return;
}
for idx in data_shards..disks.len() {
@@ -2004,7 +2066,7 @@ pub(in crate::set_disk) fn fill_deferred_bitrot_readers(
let disk = disks[idx].clone();
let data_dir = files[idx].data_dir.unwrap_or_default();
let path = format!("{object}/{data_dir}/part.{part_number}");
let reopener = copy_source_demand_bound.then(|| {
let reopener = demand_bound_lockstep.then(|| {
deferred_reader_reopener(
inline_data.clone(),
disk.clone(),
@@ -2869,6 +2931,7 @@ impl SetDisks {
read_data,
healing,
incl_free_versions,
read_data && is_get_metadata_non_inline_data_read_early_stop_enabled(),
default_parity_count,
allow_coalescing,
)
@@ -3008,6 +3071,7 @@ impl SetDisks {
read_data: bool,
healing: bool,
incl_free_versions: bool,
allow_non_inline_data_read_early_stop: bool,
default_parity_count: usize,
allow_coalescing: bool,
) -> disk::error::Result<(Vec<FileInfo>, Vec<Option<DiskError>>, MetadataFanoutDiagnostics)> {
@@ -3038,6 +3102,8 @@ impl SetDisks {
let mut scheduled_count = 0usize;
let mut force_full_wait = false;
let mut final_miss_reason_override = None;
let mut non_inline_candidate_eligible = None;
let mut single_pending_hedge_deadline = None;
let slowtail_fault = get_metadata_slowtail_fault_request(bucket.as_ref(), object.as_ref(), read_data);
let spawn_read_version =
|join_set: &mut JoinSet<(usize, disk::error::Result<FileInfo>, Duration)>, index: usize, disk: Option<DiskStore>| {
@@ -3085,17 +3151,55 @@ impl SetDisks {
}
}
while let Some(result) = join_set.join_next().await {
loop {
let mut defer_pending_inline_data_shard = false;
let result = if let Some(deadline) = single_pending_hedge_deadline.take() {
tokio::select! {
result = join_set.join_next() => result,
_ = tokio::time::sleep_until(deadline) => {
if bounded_fanout
&& !force_full_wait
&& join_set.len() == 1
&& non_inline_candidate_eligible == Some(true)
&& !accumulator.candidate_has_read_reserve()
&& next_fanout_index < disks.len()
{
while next_fanout_index < disks.len() {
let disk_index = fanout_order[next_fanout_index];
next_fanout_index = next_fanout_index.saturating_add(1);
if let Some(disk) = disks.get(disk_index).cloned() {
spawn_read_version(&mut join_set, disk_index, disk);
scheduled_count = scheduled_count.saturating_add(1);
break;
}
}
}
continue;
}
}
} else {
join_set.join_next().await
};
let Some(result) = result else { break };
match result {
Ok((index, res, elapsed)) => match res {
Ok(file_info) => {
observations.push(MetadataFanoutObservation::from_file_info(&file_info, elapsed));
accumulator.observe_file_info(&file_info);
if allow_non_inline_data_read_early_stop {
accumulator.observe_file_info_at(index, &file_info);
} else {
accumulator.observe_file_info(&file_info);
}
if allow_non_inline_data_read_early_stop && non_inline_candidate_eligible.is_none() {
non_inline_candidate_eligible =
accumulator.candidate.as_ref().map(non_inline_data_read_candidate_is_safe);
}
if bounded_fanout
&& read_data
&& !force_full_wait
&& let Some(reason) = data_read_early_stop_inline_candidate_miss_reason(&file_info)
&& !(non_inline_candidate_eligible == Some(true)
&& reason == GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_NOT_INLINE)
{
force_full_wait = true;
final_miss_reason_override.get_or_insert(reason);
@@ -3126,6 +3230,9 @@ impl SetDisks {
{
let should_return_early = if read_data {
match accumulator.candidate.as_ref() {
Some(_candidate) if non_inline_candidate_eligible == Some(true) => {
accumulator.candidate_has_read_reserve()
}
Some(candidate) => match data_read_early_stop_inline_body_miss_reason(
bucket.as_ref(),
object.as_ref(),
@@ -3196,12 +3303,37 @@ impl SetDisks {
}
let pending_responses = join_set.len();
let should_hedge_single_pending_data_read = read_data
// Inline verification can still depend on a missing data shard;
// issue one immediate spare when only that shard remains. The
// non-inline path keeps its delayed hedge below to avoid healthy
// reads paying speculative I/O before the candidate is classified.
let should_hedge_single_pending_inline_read = read_data
&& !force_full_wait
&& !defer_pending_inline_data_shard
&& pending_responses == 1
&& non_inline_candidate_eligible != Some(true)
&& accumulator.can_still_reach_early_stop_with_pending(pending_responses);
if bounded_fanout && force_full_wait {
// A non-inline plan must retain one extra matching shard as a
// reconstruction reserve. Schedule that reserve only after the
// candidate is known to be eligible, so inline GETs do not pay an
// extra fanout and the healthy path remains allocation-free.
let needs_non_inline_read_reserve = non_inline_candidate_eligible == Some(true)
&& !accumulator.candidate_has_read_reserve()
&& accumulator
.candidate_read_reserve_target()
.is_some_and(|reserve_target| scheduled_count < reserve_target || pending_responses == 0);
if bounded_fanout
&& !force_full_wait
&& (needs_non_inline_read_reserve || should_hedge_single_pending_inline_read)
&& next_fanout_index < disks.len()
{
let disk_index = fanout_order[next_fanout_index];
if let Some(disk) = disks.get(disk_index).cloned() {
spawn_read_version(&mut join_set, disk_index, disk);
scheduled_count = scheduled_count.saturating_add(1);
}
next_fanout_index = next_fanout_index.saturating_add(1);
} else if bounded_fanout && force_full_wait {
while next_fanout_index < disks.len() {
let disk_index = fanout_order[next_fanout_index];
if let Some(disk) = disks.get(disk_index).cloned() {
@@ -3213,8 +3345,7 @@ impl SetDisks {
} else if bounded_fanout
&& !defer_pending_inline_data_shard
&& next_fanout_index < disks.len()
&& (!accumulator.can_still_reach_early_stop_with_pending(pending_responses)
|| should_hedge_single_pending_data_read)
&& !accumulator.can_still_reach_early_stop_with_pending(pending_responses)
{
let disk_index = fanout_order[next_fanout_index];
if let Some(disk) = disks.get(disk_index).cloned() {
@@ -3223,6 +3354,17 @@ impl SetDisks {
}
next_fanout_index = next_fanout_index.saturating_add(1);
}
if bounded_fanout
&& !force_full_wait
&& !defer_pending_inline_data_shard
&& join_set.len() == 1
&& non_inline_candidate_eligible == Some(true)
&& !accumulator.candidate_has_read_reserve()
&& accumulator.can_still_reach_early_stop_with_pending(join_set.len())
&& next_fanout_index < disks.len()
{
single_pending_hedge_deadline = Some(tokio::time::Instant::now() + NON_INLINE_SINGLE_PENDING_HEDGE_DELAY);
}
}
let accumulator_miss_reason = accumulator.final_miss_reason();
@@ -7307,6 +7449,90 @@ mod tests {
drop(dirs);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn metadata_slowtail_fault_gate_stops_before_unneeded_tail() {
const DISKS: usize = 4;
let bucket = "metadata-slowtail-gated-bucket";
let object = "objects/metadata-slowtail-gated-object";
let (dirs, disks) = call_counter_local_disks(bucket, DISKS).await;
install_mapped_metadata_fanout_fileinfo(&disks, bucket, object).await;
let order = bounded_metadata_fanout_order(bucket, object, DISKS, 2);
let slow_disk = *order.get(3).expect("four-disk fanout should have a deferred tail disk");
let slow_disk_env = slow_disk.to_string();
temp_env::async_with_vars(
[
(ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT, Some("true")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, Some("150")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS, Some(slow_disk_env.as_str())),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET, Some(bucket)),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX, Some("objects/")),
],
async {
let calls = disk_call_counters::observe(object);
let read_with_data =
SetDisks::read_all_fileinfo_observed(&disks, bucket, bucket, object, "", true, false, false, true, 2);
let (parts_metadata, errs, diagnostics) = tokio::time::timeout(Duration::from_millis(500), read_with_data)
.await
.expect("gated metadata read should stop before the deferred slow tail")
.expect("gated metadata fanout should resolve");
assert!(parts_metadata.iter().filter(|fi| fi.name == object).count() >= 3);
assert!(errs.iter().all(Option::is_none));
assert!(diagnostics.total_responses() < DISKS);
assert_eq!(calls.total(disk_call_counters::KIND_METADATA_SLOWTAIL_FAULT), 0);
},
)
.await;
drop(dirs);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn metadata_slowtail_fault_gate_hedges_an_initial_slow_data_shard() {
const DISKS: usize = 4;
let bucket = "metadata-slowtail-gated-initial-bucket";
let object = "objects/metadata-slowtail-gated-initial-object";
let (dirs, disks) = call_counter_local_disks(bucket, DISKS).await;
install_mapped_metadata_fanout_fileinfo(&disks, bucket, object).await;
let order = bounded_metadata_fanout_order(bucket, object, DISKS, 2);
let slow_disk = *order.get(1).expect("four-disk fanout should have an initial data disk");
let spare_disk = *order.get(3).expect("four-disk fanout should have a spare disk");
let slow_disk_env = slow_disk.to_string();
temp_env::async_with_vars(
[
(ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT, Some("true")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, Some("500")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS, Some(slow_disk_env.as_str())),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET, Some(bucket)),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX, Some("objects/")),
],
async {
let calls = disk_call_counters::observe(object);
let read_with_data =
SetDisks::read_all_fileinfo_observed(&disks, bucket, bucket, object, "", true, false, false, true, 2);
let (parts_metadata, errs, diagnostics) = tokio::time::timeout(Duration::from_millis(300), read_with_data)
.await
.expect("gated metadata read should hedge the initial slow shard")
.expect("gated metadata fanout should resolve");
assert!(parts_metadata.iter().filter(|fi| fi.name == object).count() >= 3);
assert!(errs.iter().all(Option::is_none));
assert!(diagnostics.total_responses() < DISKS);
assert_eq!(calls.for_disk(disk_call_counters::KIND_METADATA_SLOWTAIL_FAULT, slow_disk), 1);
assert_eq!(calls.for_disk(disk_call_counters::KIND_READ_VERSION, spare_disk), 1);
},
)
.await;
drop(dirs);
}
/// Demo / regression guard for the backlog#1325 per-disk call counters.
///
/// The metadata fan-out issues each `read_version` inside its own
@@ -7720,6 +7946,32 @@ mod tests {
}
}
async fn install_mapped_metadata_fanout_fileinfo(disks: &[Option<DiskStore>], bucket: &str, object: &str) {
let version_id = Uuid::new_v4();
let data_dir = Uuid::new_v4();
let mod_time = OffsetDateTime::now_utc();
let distribution = FileInfo::new(&metadata_distribution_key(bucket, object), 2, 2)
.erasure
.distribution;
for (index, disk) in disks
.iter()
.enumerate()
.filter_map(|(index, disk)| disk.as_ref().map(|disk| (index, disk)))
{
disk.write_all(bucket, &format!("{object}/{data_dir}/part.1"), Bytes::from_static(b"x"))
.await
.expect("part data should be installed on every disk");
let mut file_info = valid_metadata_fanout_fileinfo(bucket, object, version_id, data_dir, mod_time);
file_info.erasure.distribution = distribution.clone();
file_info.erasure.index = *distribution
.get(index)
.expect("mapped metadata distribution should cover every disk");
disk.write_metadata(bucket, bucket, object, file_info)
.await
.expect("mapped metadata should be installed on every disk");
}
}
async fn inline_metadata_fanout_fileinfos_with_mode(
bucket: &str,
object: &str,
@@ -10342,6 +10594,41 @@ mod tests {
assert_eq!(accumulator.candidate_latest_quorum(&impossible_parity), None);
}
#[test]
fn metadata_quorum_accumulator_tracks_mapped_shards_and_requires_a_reserve() {
let version_id = Uuid::new_v4();
let data_dir = Uuid::new_v4();
let base = valid_metadata_fanout_fileinfo("bucket", "object", version_id, data_dir, OffsetDateTime::now_utc());
let distribution = base.erasure.distribution.clone();
let mut accumulator = MetadataQuorumAccumulator::new(4, 2, true);
for (disk_index, &erasure_index) in distribution.iter().take(2).enumerate() {
let mut file_info = base.clone();
file_info.erasure.index = erasure_index;
accumulator.observe_file_info_at(disk_index, &file_info);
}
assert!(
!accumulator.candidate_has_read_reserve(),
"data quorum without parity reserve must not early-stop"
);
let mut mismatched = base.clone();
mismatched.erasure.index = distribution[3];
accumulator.observe_file_info_at(2, &mismatched);
assert!(
!accumulator.candidate_has_read_reserve(),
"mapped index mismatch must not count as a reserve"
);
let mut reserve = base;
reserve.erasure.index = distribution[2];
accumulator.observe_file_info_at(2, &reserve);
assert!(
accumulator.candidate_has_read_reserve(),
"one matching reserve shard should complete the read reserve"
);
}
#[test]
fn metadata_quorum_accumulator_treats_invalid_default_parity_as_full_fanout() {
let accumulator = MetadataQuorumAccumulator::new(2, 2, true);
+361 -5
View File
@@ -773,6 +773,14 @@ const DEFAULT_RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE: bool = false;
const ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE";
const DEFAULT_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE: bool = true;
// Opt-in non-inline data-read quorum early-stop rollout (backlog#1309). The
// existing metadata fanout still reads data-bearing metadata; this gate only
// permits a safe plain single-part candidate to stop before the full fanout.
// Keep it opt-in until the Linux multi-node slow-tail and small-inline cost
// gates are complete. The environment name is retained for compatibility.
const ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE: &str = "RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE";
const DEFAULT_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE: bool = false;
const ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT: &str = "RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT";
const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT: bool = true;
@@ -1030,14 +1038,19 @@ mod prepared_get_object_metadata_tests {
const READ_VERSION_BARRIER_GUARD: std::time::Duration = std::time::Duration::from_secs(10);
fn object_with_initial_data_shards(bucket: &str, prefix: &str) -> String {
object_with_initial_data_shards_for_geometry(bucket, prefix, 4, 2)
}
fn object_with_initial_data_shards_for_geometry(bucket: &str, prefix: &str, total_disks: usize, parity: usize) -> String {
(0..1000)
.map(|index| format!("{prefix}-{index}.bin"))
.find(|name| {
let order = bounded_metadata_fanout_order(bucket, name, 4, 2);
let distribution = FileInfo::new(&[bucket, name].join("/"), 2, 2).erasure.distribution;
let mut seen = [false; 2];
for disk_index in order.into_iter().take(3) {
if let Some(block_index @ 1..=2) = distribution.get(disk_index).copied() {
let order = bounded_metadata_fanout_order(bucket, name, total_disks, parity);
let data = total_disks.saturating_sub(parity);
let distribution = FileInfo::new(&[bucket, name].join("/"), data, parity).erasure.distribution;
let mut seen = vec![false; data];
for disk_index in order.into_iter().take(total_disks.saturating_sub(parity).saturating_add(1)) {
if let Some(block_index) = distribution.get(disk_index).copied().filter(|index| *index <= data) {
seen[block_index - 1] = true;
}
}
@@ -1168,6 +1181,223 @@ mod prepared_get_object_metadata_tests {
);
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn non_inline_data_read_early_stop_uses_quorum_plan() {
let (_dirs, set_disks) = make_local_set_disks(4, 2).await;
let bucket = "non-inline-read-plan";
let object = object_with_initial_data_shards(bucket, "non-inline-object");
let payload = vec![0x5a; 2 * 1024 * 1024];
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut put_reader = PutObjReader::from_vec(payload.clone());
set_disks
.put_object(bucket, &object, &mut put_reader, &opts)
.await
.expect("object should be written");
temp_env::async_with_vars(
[
("RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")),
],
async {
let calls = disk_call_counters::observe(&object);
reset_test_get_object_reader_path();
let mut reader = set_disks
.get_object_reader(bucket, &object, None, HeaderMap::new(), &opts)
.await
.expect("quorum GET reader should open");
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("quorum GET body should stream");
assert_eq!(restored, payload);
assert!(
calls.total(disk_call_counters::KIND_READ_VERSION) < 4,
"non-inline quorum GET should retain a reconstruction reserve"
);
},
)
.await;
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn non_inline_data_read_early_stop_keeps_reserve_on_unequal_layout() {
let (_dirs, set_disks) = make_local_set_disks(6, 2).await;
let bucket = "non-inline-read-reserve";
let object = object_with_initial_data_shards_for_geometry(bucket, "reserve-object", 6, 2);
let payload = vec![0x5a; 2 * 1024 * 1024];
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut put_reader = PutObjReader::from_vec(payload.clone());
set_disks
.put_object(bucket, &object, &mut put_reader, &opts)
.await
.expect("object should be written");
temp_env::async_with_vars(
[
("RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")),
],
async {
let calls = disk_call_counters::observe(&object);
let mut reader = set_disks
.get_object_reader(bucket, &object, None, HeaderMap::new(), &opts)
.await
.expect("quorum GET reader should open");
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("quorum GET body should stream");
assert_eq!(restored, payload);
assert_eq!(
calls.total(disk_call_counters::KIND_READ_VERSION),
5,
"the unequal layout should schedule exactly one reserve beyond its data quorum"
);
},
)
.await;
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn non_inline_data_read_early_stop_preserves_inline_path() {
let (_dirs, set_disks) = make_local_set_disks(4, 2).await;
let bucket = "non-inline-read-plan-inline";
let object = object_with_initial_data_shards(bucket, "inline-object");
let payload = b"quorum inline payload".repeat(256);
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut put_reader = PutObjReader::from_vec(payload.clone());
set_disks
.put_object(bucket, &object, &mut put_reader, &opts)
.await
.expect("inline object should be written");
temp_env::async_with_vars(
[
("RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")),
],
async {
let calls = disk_call_counters::observe(&object);
let mut reader = set_disks
.get_object_reader(bucket, &object, None, HeaderMap::new(), &opts)
.await
.expect("inline GET reader should open");
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("inline GET body should stream");
assert_eq!(restored, payload);
assert_eq!(
test_get_object_reader_path_id(),
3,
"inline GET should retain the direct inline reader path"
);
assert!(calls.total(disk_call_counters::KIND_READ_VERSION) <= 4);
},
)
.await;
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn non_inline_data_read_early_stop_does_not_add_inline_fanout_on_unequal_layout() {
let (_dirs, set_disks) = make_local_set_disks(6, 2).await;
let bucket = "inline-read-plan-unequal";
let object = object_with_initial_data_shards_for_geometry(bucket, "inline-object", 6, 2);
let payload = b"inline quorum payload".repeat(256);
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut put_reader = PutObjReader::from_vec(payload.clone());
set_disks
.put_object(bucket, &object, &mut put_reader, &opts)
.await
.expect("inline object should be written");
let read_once = |enabled: bool| {
let set_disks = Arc::clone(&set_disks);
let bucket = bucket.to_string();
let object = object.clone();
let payload = payload.clone();
let opts = opts.clone();
async move {
temp_env::async_with_vars(
[
(
"RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE",
Some(if enabled { "true" } else { "false" }),
),
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")),
],
async {
let calls = disk_call_counters::observe(&object);
let mut reader = set_disks
.get_object_reader(&bucket, &object, None, HeaderMap::new(), &opts)
.await
.expect("inline GET reader should open");
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("inline GET body should stream");
assert_eq!(restored, payload);
calls.total(disk_call_counters::KIND_READ_VERSION)
},
)
.await
}
};
let gate_off_calls = read_once(false).await;
let gate_on_calls = read_once(true).await;
assert_eq!(gate_on_calls, gate_off_calls, "inline gate must not add reserve fanout");
}
#[test]
#[serial_test::serial(body_cache_hook)]
fn inline_data_read_early_stop_defaults_return_exact_body() {
@@ -1914,6 +2144,26 @@ fn is_get_metadata_data_read_early_stop_enabled() -> bool {
}
}
fn is_get_metadata_non_inline_data_read_early_stop_enabled() -> bool {
#[cfg(test)]
{
rustfs_utils::get_env_bool(
ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE,
DEFAULT_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE,
)
}
#[cfg(not(test))]
{
static CACHED: OnceLock<bool> = OnceLock::new();
*CACHED.get_or_init(|| {
rustfs_utils::get_env_bool(
ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE,
DEFAULT_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE,
)
})
}
}
fn is_get_metadata_early_stop_bounded_fanout_enabled() -> bool {
#[cfg(test)]
{
@@ -6241,6 +6491,7 @@ mod tests {
use crate::object_api::BLOCK_SIZE_V2;
use crate::object_api::ObjectInfo;
use crate::set_disk::core::io_primitives::rename_fanout_barrier;
use crate::set_disk::ops::object::{PutObjectCommitBarrier, PutObjectCommitPause};
use crate::storage_api_contracts::{
heal::HealOperations as _, lifecycle::TransitionedObject, list::ListOperations as _, multipart::CompletePart,
object::ObjectOperations as _,
@@ -12681,6 +12932,111 @@ mod tests {
.await;
}
#[tokio::test(flavor = "multi_thread")]
#[serial]
async fn multipart_streaming_get_blocks_overwrite_across_part_boundary() {
temp_env::async_with_vars(
[
(rustfs_config::ENV_OBJECT_LOCK_OPTIMIZATION_ENABLE, Some("true")),
(ENV_RUSTFS_GET_MULTIPART_READER_SETUP_PREFETCH, Some("false")),
],
async {
let set_disks = make_local_bucket_test_set_disks().await;
let bucket = "snapshot-multipart-overwrite";
let object = "object";
let part_size = usize::try_from(GLOBAL_MIN_PART_SIZE.as_u64()).expect("minimum part size should fit usize");
let first_part = vec![0x41; part_size];
let second_part = vec![0x42; part_size];
let replacement = vec![0x43; first_part.len() + second_part.len()];
let opts = ObjectOptions::default();
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let upload = set_disks
.new_multipart_upload(bucket, object, &opts)
.await
.expect("multipart upload should be created");
let mut completed_parts = Vec::with_capacity(2);
for (part_num, body) in [(1, &first_part), (2, &second_part)] {
let mut reader = PutObjReader::from_vec(body.clone());
let part = set_disks
.put_object_part(bucket, object, &upload.upload_id, part_num, &mut reader, &opts)
.await
.expect("multipart part should be written");
completed_parts.push(CompletePart {
part_num,
etag: part.etag,
..Default::default()
});
}
let completed = Arc::clone(&set_disks)
.complete_multipart_upload(bucket, object, &upload.upload_id, completed_parts, &opts)
.await
.expect("multipart upload should complete");
assert!(completed.is_multipart());
let mut snapshot = set_disks
.get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
.await
.expect("multipart snapshot reader should open");
let overwrite_set = Arc::clone(&set_disks);
let overwrite_opts = opts.clone();
let overwrite_body = replacement.clone();
let commit_barrier = PutObjectCommitBarrier::install(bucket, object, PutObjectCommitPause::BeforeNamespace);
let overwrite = tokio::spawn(async move {
let mut reader = PutObjReader::from_vec(overwrite_body);
overwrite_set.put_object(bucket, object, &mut reader, &overwrite_opts).await
});
commit_barrier.wait_until_paused().await;
commit_barrier.release_and_wait_until_namespace_pending().await;
assert!(
!commit_barrier.namespace_acquired(),
"overwrite must wait for the multipart response's read lock"
);
let mut restored_first = vec![0; first_part.len()];
snapshot
.stream
.read_exact(&mut restored_first)
.await
.expect("the first multipart part should stream");
assert_eq!(restored_first, first_part);
assert!(
!commit_barrier.namespace_acquired() && !overwrite.is_finished(),
"overwrite must remain blocked at the first/second part boundary"
);
let mut restored_second = Vec::new();
snapshot
.stream
.read_to_end(&mut restored_second)
.await
.expect("the second multipart part should stream");
assert_eq!(restored_second, second_part);
tokio::time::timeout(Duration::from_secs(5), overwrite)
.await
.expect("overwrite should proceed after multipart EOF")
.expect("overwrite task should join")
.expect("overwrite should succeed");
let mut latest = set_disks
.get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
.await
.expect("replacement reader should open");
let mut latest_body = Vec::new();
latest
.stream
.read_to_end(&mut latest_body)
.await
.expect("replacement should stream");
assert_eq!(latest_body, replacement);
},
)
.await;
}
#[tokio::test(flavor = "multi_thread")]
#[serial]
async fn streaming_get_blocks_concurrent_delete_until_eof() {
+17 -11
View File
@@ -1675,7 +1675,7 @@ impl crate::storage_api_contracts::object::ObjectIO for SetDisks {
None
};
let metadata_stage_start = Instant::now();
let metadata_stage_start = stage_metrics_enabled.then(Instant::now);
let (snapshot, prepared_object_info) = if let Some(prepared) = take_prepared_get_object_metadata() {
(prepared.snapshot, prepared.object_info)
} else {
@@ -1691,7 +1691,11 @@ impl crate::storage_api_contracts::object::ObjectIO for SetDisks {
{
Ok(snapshot) => (snapshot, None),
Err(err) => {
rustfs_io_metrics::record_get_object_metadata_phase_duration(metadata_stage_start.elapsed().as_secs_f64());
if let Some(metadata_stage_start) = metadata_stage_start {
rustfs_io_metrics::record_get_object_metadata_phase_duration(
metadata_stage_start.elapsed().as_secs_f64(),
);
}
let failure_path = if is_meta_bucketname(bucket) {
GET_OBJECT_PATH_INTERNAL_META
} else {
@@ -1716,15 +1720,17 @@ impl crate::storage_api_contracts::object::ObjectIO for SetDisks {
};
let size_bucket = rustfs_io_metrics::get_object_size_bucket(metrics_size);
record_get_stage_duration_if_enabled(GET_OBJECT_PATH_SET_DISK, GET_STAGE_OBJECT_INFO, object_info_stage_start);
let metadata_elapsed = metadata_stage_start.elapsed().as_secs_f64();
rustfs_io_metrics::record_get_object_metadata_phase_duration(metadata_elapsed);
rustfs_io_metrics::record_get_object_stage_duration_by_size(
GET_OBJECT_PATH_SET_DISK,
GET_STAGE_METADATA,
object_class.as_str(),
size_bucket,
metadata_elapsed,
);
if let Some(metadata_stage_start) = metadata_stage_start {
let metadata_elapsed = metadata_stage_start.elapsed().as_secs_f64();
rustfs_io_metrics::record_get_object_metadata_phase_duration(metadata_elapsed);
rustfs_io_metrics::record_get_object_stage_duration_by_size(
GET_OBJECT_PATH_SET_DISK,
GET_STAGE_METADATA,
object_class.as_str(),
size_bucket,
metadata_elapsed,
);
}
if object_info.delete_marker {
if opts.version_id.is_none() {
+39 -9
View File
@@ -89,6 +89,8 @@ use super::ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT;
#[cfg(test)]
use super::ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE;
#[cfg(test)]
use super::ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE;
#[cfg(test)]
use super::ENV_RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE;
#[cfg(test)]
use super::ENV_RUSTFS_GET_MULTIPART_READER_SETUP_PREFETCH;
@@ -126,6 +128,8 @@ use super::is_get_metadata_early_stop_bounded_fanout_enabled;
#[cfg(test)]
use super::is_get_metadata_early_stop_enabled;
#[cfg(test)]
use super::is_get_metadata_non_inline_data_read_early_stop_enabled;
#[cfg(test)]
use super::is_version_early_stop_enabled;
#[cfg(test)]
use super::load_get_codec_streaming_config;
@@ -4267,6 +4271,20 @@ mod tests {
);
}
#[test]
#[serial(body_cache_hook)]
fn non_inline_data_read_early_stop_gate_defaults_off_and_honors_override() {
temp_env::with_var(ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE, None::<&str>, || {
assert!(!is_get_metadata_non_inline_data_read_early_stop_enabled());
});
temp_env::with_var(ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE, Some("true"), || {
assert!(is_get_metadata_non_inline_data_read_early_stop_enabled());
});
temp_env::with_var(ENV_RUSTFS_GET_METADATA_TWO_PHASE_READ_PLAN_ENABLE, Some("false"), || {
assert!(!is_get_metadata_non_inline_data_read_early_stop_enabled());
});
}
#[test]
fn metadata_early_stop_rejects_healing_and_free_version_requests() {
temp_env::with_vars(
@@ -5547,9 +5565,10 @@ mod tests {
/// backlog#923: with the data-shards-only lockstep gate on, every retained
/// parity reader must be an unopened deferred reader carrying a stripe
/// handle, so the decode path can realign it to a mid-object stripe. With
/// the gate off (default), eagerly opened parity readers are kept exactly
/// as before and carry no handles.
/// handle and disposable reopener, so the decode path can realign it to a
/// mid-object stripe without consuming the later-stripe reserve. With the
/// gate off (default), eagerly opened parity readers are kept exactly as
/// before and carry neither.
#[tokio::test]
#[serial_test::serial]
async fn bitrot_reader_setup_gates_parity_stripe_handle_conversion() {
@@ -5578,6 +5597,11 @@ mod tests {
enabled.is_some(),
"parity slot {idx} stripe handle must match the gate (enabled={enabled:?})"
);
assert_eq!(
setup.deferred_reopeners[idx].is_some(),
enabled.is_some(),
"parity slot {idx} reopener must match the gate (enabled={enabled:?})"
);
}
if enabled.is_some() {
@@ -5600,13 +5624,17 @@ mod tests {
}
#[tokio::test]
#[serial_test::serial]
async fn bitrot_reader_setup_data_blocks_first_keeps_deferred_fallback_readers() {
let mut setup = setup_inline_bitrot_readers_with_env(
vec![Some(b"aaaa"), Some(b"bbbb"), Some(b"cccc"), Some(b"dddd")],
2,
2,
BitrotReaderSetupMode::ReadQuorum,
true,
let mut setup = temp_env::async_with_vars(
[("RUSTFS_GET_LOCKSTEP_DATA_SHARDS_ONLY_ENABLE", Some("true"))],
setup_inline_bitrot_readers_with_env(
vec![Some(b"aaaa"), Some(b"bbbb"), Some(b"cccc"), Some(b"dddd")],
2,
2,
BitrotReaderSetupMode::ReadQuorum,
true,
),
)
.await;
@@ -5614,6 +5642,8 @@ mod tests {
assert_eq!(setup.available_shards(), 2);
assert_eq!(setup.scheduled_shards(), 2);
assert_eq!(setup.readers.iter().filter(|reader| reader.is_some()).count(), 4);
assert!(setup.deferred_reopeners[2].is_some());
assert!(setup.deferred_reopeners[3].is_some());
let fallback_index = setup
.attempted
+355 -25
View File
@@ -1817,16 +1817,21 @@ impl ECStore {
let metadata = pool.prepare_get_object_reader_metadata(bucket, &object, &opts).await?;
(metadata, pool)
} else {
let (_, pool_idx) = self
.get_latest_accessible_object_info_with_idx(bucket, &object, &opts)
.await?;
let pool = self
.pools
.get(pool_idx)
.cloned()
.ok_or_else(|| Error::other(format!("resolved GET pool index {pool_idx} is out of bounds")))?;
let metadata = pool.prepare_get_object_reader_metadata(bucket, &object, &opts).await?;
(metadata, pool)
// Keep the large multi-pool selection future off the caller stack.
// Debug builds otherwise exceed the common 2 MiB worker stack.
Box::pin(async {
let (metadata, pool_idx) = self.prepare_latest_object_metadata_with_idx(bucket, &object, &opts).await?;
if let Some(error) = latest_object_access_delete_marker_error(bucket, &object, metadata.object_info(), &opts) {
return Err(error);
}
let pool = self
.pools
.get(pool_idx)
.cloned()
.ok_or_else(|| Error::other(format!("resolved GET pool index {pool_idx} is out of bounds")))?;
Ok((metadata, pool))
})
.await?
};
Ok(PreparedGetObjectReader {
@@ -2518,12 +2523,18 @@ impl ECStore {
.get_object_reader(bucket, object.as_ref(), range, h, &opts)
.await?
} else {
let (_, idx) = self
.get_latest_accessible_object_info_with_idx(bucket, &object, &opts)
.await?;
self.pools[idx]
.get_object_reader(bucket, object.as_ref(), range, h, &opts)
.await?
// Keep selection plus prepared-open state off the caller stack.
// Debug builds otherwise exceed the common 2 MiB worker stack.
Box::pin(async {
let (metadata, idx) = self.prepare_latest_object_metadata_with_idx(bucket, &object, &opts).await?;
if let Some(error) = latest_object_access_delete_marker_error(bucket, &object, metadata.object_info(), &opts) {
return Err(error);
}
self.pools[idx]
.get_object_reader_with_prepared_metadata(bucket, object.as_ref(), range, h, &opts, metadata)
.await
})
.await?
};
Ok(Self::attach_read_lock_guard(reader, read_lock_guard))
@@ -3914,6 +3925,7 @@ mod tests {
ReplicationState, ReplicationStatusType, VersionPurgeStatusType, replication_state_to_filemeta, replication_statuses_map,
version_purge_statuses_map,
};
use crate::core::pools::{PoolDecommissionInfo, PoolStatus};
use crate::core::sets::make_local_two_set_sets_with_ctx;
use crate::ecstore_validation_blackbox::{make_local_set_disks, make_local_set_disks_with_ctx};
use crate::layout::{
@@ -6085,13 +6097,62 @@ mod tests {
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn prepared_reader_resolves_object_from_second_pool() {
async fn prepared_reader_reuses_metadata_across_three_pools() {
let (_first_dirs, first_set) = make_local_set_disks(4, 2).await;
let (_second_dirs, second_set) = make_local_set_disks(4, 2).await;
let (_third_dirs, third_set) = make_local_set_disks(4, 2).await;
let store = new_prepared_reader_test_store(&[first_set, second_set, third_set]).await;
let bucket = "prepared-reader-three-pools";
let object = "object.bin";
let payload = b"prepared-reader-three-pool-payload-".repeat(40_000);
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
for pool in &store.pools {
pool.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created in each pool");
}
let mut put_reader = PutObjReader::from_vec(payload.clone());
store.pools[2]
.put_object(bucket, object, &mut put_reader, &opts)
.await
.expect("object should be written only to the third pool");
let calls = disk_call_counters::observe(object);
let prepared = store
.prepare_get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
.await
.expect("prepared reader should resolve the third-pool object");
assert_eq!(prepared.object_info().size, payload.len() as i64);
let metadata_calls = calls.total(disk_call_counters::KIND_READ_VERSION);
assert_eq!(metadata_calls, 12, "three 4-disk pools must fan out metadata exactly once each");
let mut reader = prepared.into_reader().await.expect("prepared body reader should open");
assert_eq!(
calls.total(disk_call_counters::KIND_READ_VERSION),
metadata_calls,
"the selected pool must reuse its prepared metadata"
);
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("prepared body should stream");
assert_eq!(restored, payload);
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn legacy_reader_reuses_selected_pool_metadata() {
let (_first_dirs, first_set) = make_local_set_disks(4, 2).await;
let (_second_dirs, second_set) = make_local_set_disks(4, 2).await;
let store = new_prepared_reader_test_store(&[first_set, second_set]).await;
let bucket = "prepared-reader-second-pool";
let bucket = "legacy-reader-second-pool";
let object = "object.bin";
let payload = b"prepared-reader-second-pool-payload-".repeat(40_000);
let payload = b"legacy-reader-second-pool-payload-".repeat(40_000);
let opts = ObjectOptions {
no_lock: true,
..Default::default()
@@ -6108,18 +6169,287 @@ mod tests {
.await
.expect("object should be written only to the second pool");
let prepared = store
.prepare_get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
clear_get_object_body_cache_hook();
let hook = Arc::new(CountingMissHook {
calls: AtomicUsize::new(0),
});
register_get_object_body_cache_hook(Arc::clone(&hook) as Arc<dyn GetObjectBodyCacheHook>);
let _hook_guard = BodyCacheHookGuard;
let calls = disk_call_counters::observe(object);
let mut reader = store
.handle_get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
.await
.expect("prepared reader should resolve the second-pool object");
assert_eq!(prepared.object_info().size, payload.len() as i64);
let mut reader = prepared.into_reader().await.expect("prepared body reader should open");
.expect("legacy reader should resolve the second-pool object");
assert_eq!(
hook.calls.load(Ordering::Relaxed),
1,
"legacy reader must probe the body cache exactly once"
);
assert_eq!(reader.body_source, GetObjectBodySource::HookMissed);
assert!(
calls.total(disk_call_counters::KIND_READ_VERSION) <= 8,
"legacy reader must fan out each 4-disk pool at most once"
);
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("prepared body should stream");
.expect("legacy reader body should stream");
assert_eq!(restored, payload);
}
fn prepared_pool_test_status(id: usize, suspended: bool) -> PoolStatus {
PoolStatus {
id,
cmd_line: format!("prepared-pool-{id}"),
last_update: OffsetDateTime::now_utc(),
decommission: suspended.then(|| PoolDecommissionInfo {
start_time: Some(OffsetDateTime::now_utc()),
..Default::default()
}),
}
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn prepared_reader_refetches_when_final_pool_state_changes_winner() {
let (_dirs, set_disks) = make_local_set_disks(4, 2).await;
let store = Arc::new(new_prepared_reader_test_store(&[Arc::clone(&set_disks), Arc::clone(&set_disks)]).await);
let bucket = "prepared-reader-pool-state-fallback";
let object = "object.bin";
let payload = b"pool-state fallback payload".repeat(8_000);
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut put_reader = PutObjReader::from_vec(payload.clone());
set_disks
.put_object(bucket, object, &mut put_reader, &opts)
.await
.expect("shared object should be written");
let calls = disk_call_counters::observe(object);
let barrier = crate::store::rebalance::PreparedPoolReadFallbackBarrier::install(object, false);
let read_store = Arc::clone(&store);
let read_opts = opts.clone();
let read = tokio::spawn(async move {
read_store
.prepare_get_object_reader(bucket, object, None, HeaderMap::new(), &read_opts)
.await
});
barrier.wait_after_fanout().await;
*store.pool_meta.write().await = PoolMeta {
pools: vec![prepared_pool_test_status(0, false), prepared_pool_test_status(1, true)],
..Default::default()
};
barrier.release_after_fanout();
let prepared = read
.await
.expect("prepared read task should not panic")
.expect("final active pool should be refetched");
assert!(Arc::ptr_eq(&prepared.pool, &store.pools[0]));
assert_eq!(
calls.total(disk_call_counters::KIND_READ_VERSION),
12,
"two initial 4-disk fanouts plus one fallback refetch are required"
);
let mut reader = prepared.into_reader().await.expect("fallback body reader should open");
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("fallback body should stream");
assert_eq!(restored, payload);
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn prepared_reader_fallback_rejects_generation_change_before_refetch() {
let (_dirs, set_disks) = make_local_set_disks(4, 2).await;
let store = Arc::new(new_prepared_reader_test_store(&[Arc::clone(&set_disks), Arc::clone(&set_disks)]).await);
let bucket = "prepared-reader-pool-state-generation-change";
let object = "object.bin";
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut initial_reader = PutObjReader::from_vec(b"initial generation".to_vec());
set_disks
.put_object(bucket, object, &mut initial_reader, &opts)
.await
.expect("initial object should be written");
let barrier = crate::store::rebalance::PreparedPoolReadFallbackBarrier::install(object, true);
let read_store = Arc::clone(&store);
let read_opts = opts.clone();
let read = tokio::spawn(async move {
read_store
.prepare_get_object_reader(bucket, object, None, HeaderMap::new(), &read_opts)
.await
});
barrier.wait_after_fanout().await;
*store.pool_meta.write().await = PoolMeta {
pools: vec![prepared_pool_test_status(0, false), prepared_pool_test_status(1, true)],
..Default::default()
};
barrier.release_after_fanout();
barrier.wait_before_refetch().await;
let mut replacement_reader = PutObjReader::from_vec(b"replacement generation".to_vec());
set_disks
.put_object(bucket, object, &mut replacement_reader, &opts)
.await
.expect("replacement generation should be written before fallback refetch");
barrier.release_before_refetch();
let error = match read.await.expect("prepared read task should not panic") {
Ok(_) => panic!("changed fallback generation must not be accepted"),
Err(error) => error,
};
assert_eq!(error, Error::ErasureReadQuorum);
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn prepared_reader_rejects_latest_delete_marker_without_refetching_metadata() {
let ctx = Arc::new(crate::runtime::instance::InstanceContext::new());
let (_first_dirs, first_set) = make_local_set_disks_with_ctx(4, 2, Arc::clone(&ctx)).await;
let (_second_dirs, second_set) = make_local_set_disks_with_ctx(4, 2, Arc::clone(&ctx)).await;
let store = new_prepared_reader_test_store_with_ctx(&[Arc::clone(&first_set), Arc::clone(&second_set)], ctx).await;
let bucket = "prepared-reader-latest-delete-marker";
let object = "versioned-object.bin";
let versioned_opts = ObjectOptions {
no_lock: true,
versioned: true,
object_lock_config_snapshot: Some(Arc::new(ObjectLockConfigSnapshot::new(ObjectLockConfigState::ConfirmedAbsent))),
..Default::default()
};
for set_disks in [&first_set, &second_set] {
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
}
let mut older_reader = PutObjReader::from_vec(b"older visible generation".to_vec());
first_set
.put_object(bucket, object, &mut older_reader, &versioned_opts)
.await
.expect("older object should be written");
let mut hidden_reader = PutObjReader::from_vec(b"hidden generation".to_vec());
second_set
.put_object(bucket, object, &mut hidden_reader, &versioned_opts)
.await
.expect("newer object should be written");
let marker = second_set
.delete_object(bucket, object, versioned_opts.clone())
.await
.expect("delete marker should be committed");
assert!(marker.delete_marker);
let calls = disk_call_counters::observe(object);
let error = match store
.prepare_get_object_reader(
bucket,
object,
None,
HeaderMap::new(),
&ObjectOptions {
no_lock: true,
versioned: true,
..Default::default()
},
)
.await
{
Ok(_) => panic!("latest delete marker should hide the older live object"),
Err(error) => error,
};
assert!(is_err_object_not_found(&error));
assert!(
calls.total(disk_call_counters::KIND_READ_VERSION) <= 8,
"delete-marker resolution must fan out each pool at most once"
);
}
#[tokio::test]
#[serial_test::serial(body_cache_hook)]
async fn prepared_reader_explicit_version_reuses_the_matching_pool_metadata() {
let (_first_dirs, first_set) = make_local_set_disks(4, 2).await;
let (_second_dirs, second_set) = make_local_set_disks(4, 2).await;
let store = new_prepared_reader_test_store(&[Arc::clone(&first_set), Arc::clone(&second_set)]).await;
let bucket = "prepared-reader-explicit-version";
let object = "versioned-object.bin";
let payload = b"explicit version from first pool".repeat(8_000);
let versioned_opts = ObjectOptions {
no_lock: true,
versioned: true,
object_lock_config_snapshot: Some(Arc::new(ObjectLockConfigSnapshot::new(ObjectLockConfigState::ConfirmedAbsent))),
..Default::default()
};
for set_disks in [&first_set, &second_set] {
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
}
let mut first_reader = PutObjReader::from_vec(payload.clone());
let first = first_set
.put_object(bucket, object, &mut first_reader, &versioned_opts)
.await
.expect("requested version should be written to the first pool");
let mut second_reader = PutObjReader::from_vec(b"different pool version".to_vec());
second_set
.put_object(bucket, object, &mut second_reader, &versioned_opts)
.await
.expect("a different version should be written to the second pool");
let requested_version = first
.version_id
.expect("versioned PUT should return a version id")
.to_string();
let read_opts = ObjectOptions {
no_lock: true,
versioned: true,
version_id: Some(requested_version),
..Default::default()
};
let calls = disk_call_counters::observe(object);
let prepared = store
.prepare_get_object_reader(bucket, object, None, HeaderMap::new(), &read_opts)
.await
.expect("explicit version should resolve from the matching pool");
assert_eq!(prepared.object_info().version_id, first.version_id);
let metadata_calls = calls.total(disk_call_counters::KIND_READ_VERSION);
assert!(metadata_calls <= 8, "explicit-version lookup must fan out each pool at most once");
let mut reader = prepared
.into_reader()
.await
.expect("prepared explicit-version body should open");
assert_eq!(calls.total(disk_call_counters::KIND_READ_VERSION), metadata_calls);
let mut restored = Vec::new();
reader
.stream
.read_to_end(&mut restored)
.await
.expect("explicit-version body should stream");
assert_eq!(restored, payload);
}
+228 -1
View File
@@ -18,18 +18,117 @@ use crate::core::pools::merge_pool_status_refresh;
use crate::layout::pool_space::{ServerPoolsAvailableSpace, build_server_pools_available_space};
use crate::runtime::sources as runtime_sources;
use crate::storage_api_contracts::{admin::StorageAdminApi, namespace::NamespaceLocking as _, object::ObjectOperations as _};
use futures::stream::{FuturesUnordered, StreamExt};
pub(in crate::store) mod support;
const LOG_COMPONENT_ECSTORE: &str = "ecstore";
const LOG_SUBSYSTEM_POOLS: &str = "pools";
const EVENT_POOL_META_RELOAD: &str = "pool_meta_reload";
#[cfg(test)]
struct PreparedPoolReadFallbackBarrierState {
object: String,
pause_before_refetch: bool,
fanout_arrived: tokio::sync::Notify,
fanout_release: tokio::sync::Notify,
refetch_arrived: tokio::sync::Notify,
refetch_release: tokio::sync::Notify,
}
#[cfg(test)]
pub(in crate::store) struct PreparedPoolReadFallbackBarrier {
state: Arc<PreparedPoolReadFallbackBarrierState>,
}
#[cfg(test)]
static PREPARED_POOL_READ_FALLBACK_BARRIER: std::sync::OnceLock<
std::sync::Mutex<Option<Arc<PreparedPoolReadFallbackBarrierState>>>,
> = std::sync::OnceLock::new();
#[cfg(test)]
impl PreparedPoolReadFallbackBarrier {
pub(in crate::store) fn install(object: &str, pause_before_refetch: bool) -> Self {
let state = Arc::new(PreparedPoolReadFallbackBarrierState {
object: object.to_string(),
pause_before_refetch,
fanout_arrived: tokio::sync::Notify::new(),
fanout_release: tokio::sync::Notify::new(),
refetch_arrived: tokio::sync::Notify::new(),
refetch_release: tokio::sync::Notify::new(),
});
*PREPARED_POOL_READ_FALLBACK_BARRIER
.get_or_init(|| std::sync::Mutex::new(None))
.lock()
.expect("prepared pool read fallback barrier must not be poisoned") = Some(Arc::clone(&state));
Self { state }
}
pub(in crate::store) async fn wait_after_fanout(&self) {
self.state.fanout_arrived.notified().await;
}
pub(in crate::store) fn release_after_fanout(&self) {
self.state.fanout_release.notify_one();
}
pub(in crate::store) async fn wait_before_refetch(&self) {
self.state.refetch_arrived.notified().await;
}
pub(in crate::store) fn release_before_refetch(&self) {
self.state.refetch_release.notify_one();
}
}
#[cfg(test)]
impl Drop for PreparedPoolReadFallbackBarrier {
fn drop(&mut self) {
self.state.fanout_release.notify_waiters();
self.state.refetch_release.notify_waiters();
if let Some(barrier) = PREPARED_POOL_READ_FALLBACK_BARRIER.get() {
*barrier
.lock()
.expect("prepared pool read fallback barrier must not be poisoned") = None;
}
}
}
#[cfg(test)]
async fn pause_prepared_pool_read_after_fanout(object: &str) {
let state = PREPARED_POOL_READ_FALLBACK_BARRIER
.get_or_init(|| std::sync::Mutex::new(None))
.lock()
.expect("prepared pool read fallback barrier must not be poisoned")
.as_ref()
.filter(|state| state.object == object)
.cloned();
if let Some(state) = state {
state.fanout_arrived.notify_one();
state.fanout_release.notified().await;
}
}
#[cfg(test)]
async fn pause_prepared_pool_read_before_refetch(object: &str) {
let state = PREPARED_POOL_READ_FALLBACK_BARRIER
.get_or_init(|| std::sync::Mutex::new(None))
.lock()
.expect("prepared pool read fallback barrier must not be poisoned")
.as_ref()
.filter(|state| state.object == object && state.pause_before_refetch)
.cloned();
if let Some(state) = state {
state.refetch_arrived.notify_one();
state.refetch_release.notified().await;
}
}
#[cfg(test)]
use support::resolve_latest_object_info_candidates;
use support::{
LatestObjectInfoCandidate, PoolErr, PoolObjInfo, RebalanceDeletePoolResult, pool_lookup_not_found_error,
rebalance_disk_set_lookup_error, resolve_latest_object_info_candidates_with_pool_state,
resolve_rebalance_delete_from_all_pools_result, resolve_rebalance_delete_from_all_pools_results,
resolve_store_rebalance_pool_meta_reload_result,
resolve_store_rebalance_pool_meta_reload_result, validate_prepared_pool_refetch_identity,
};
#[derive(Debug, Default, Eq, PartialEq)]
@@ -675,6 +774,134 @@ impl ECStore {
resolve_latest_object_info_candidates_with_pool_state(candidates, &suspended_pools, bucket, object, opts)
}
pub(super) async fn prepare_latest_object_metadata_with_idx(
&self,
bucket: &str,
object: &str,
opts: &ObjectOptions,
) -> Result<(crate::set_disk::PreparedGetObjectMetadata, usize)> {
let suspended_pools = if opts.skip_decommissioned {
let pool_meta = self.pool_meta.read().await;
Some(
(0..self.pools.len())
.map(|idx| pool_meta.is_suspended(idx))
.collect::<Vec<_>>(),
)
} else {
None
};
let mut futures = FuturesUnordered::new();
for (idx, pool) in self.pools.iter().enumerate() {
if suspended_pools.as_ref().is_some_and(|pools| pools[idx]) {
continue;
}
if opts.skip_rebalancing && self.is_pool_rebalancing(idx).await {
continue;
}
futures.push(async move {
let result = pool
.prepare_get_object_reader_metadata(bucket, object, opts)
.await
.map_err(|err| to_object_err(err, vec![bucket, object]));
(idx, result)
});
}
let mut candidates = (0..self.pools.len()).map(|_| None).collect::<Vec<_>>();
// Retain one provisional winner. Other pools only need their lightweight
// identity for final conflict checks; if pool state changes while the
// fanout runs, the final winner is refetched and revalidated below.
let mut latest_prepared = None;
let mut latest_mod_time = None;
let mut provisional_dynamic_pool_state = None;
while let Some((idx, result)) = futures.next().await {
match result {
Ok(metadata) => {
let mod_time = metadata.object_info().mod_time.unwrap_or(OffsetDateTime::UNIX_EPOCH);
let info = metadata.object_info().clone();
let retain = match (latest_mod_time, latest_prepared.as_ref()) {
(None, _) => true,
(Some(current), _) if mod_time > current => true,
(Some(current), _) if mod_time < current => false,
(Some(_), Some((current_idx, _))) => {
if suspended_pools.is_none() && provisional_dynamic_pool_state.is_none() {
let pool_meta = self.pool_meta.read().await;
provisional_dynamic_pool_state = Some(
(0..self.pools.len())
.map(|pool_idx| pool_meta.is_suspended(pool_idx))
.collect::<Vec<_>>(),
);
}
let provisional_pool_state = suspended_pools
.as_ref()
.or(provisional_dynamic_pool_state.as_ref())
.ok_or_else(|| Error::other("GET pool state snapshot is unavailable"))?;
let new_key = (provisional_pool_state.get(idx).copied().unwrap_or(false), std::cmp::Reverse(idx));
let current_key = (
provisional_pool_state.get(*current_idx).copied().unwrap_or(false),
std::cmp::Reverse(*current_idx),
);
new_key < current_key
}
(Some(_), None) => true,
};
if retain {
if latest_mod_time.is_none_or(|current| mod_time > current) {
latest_mod_time = Some(mod_time);
}
latest_prepared = Some((idx, metadata));
}
candidates[idx] = Some(LatestObjectInfoCandidate {
info: Some(info),
idx,
err: None,
});
}
Err(err) => {
candidates[idx] = Some(LatestObjectInfoCandidate {
info: None,
idx,
err: Some(err),
});
}
}
}
#[cfg(test)]
pause_prepared_pool_read_after_fanout(object).await;
let suspended_pools = match suspended_pools {
Some(pools) => pools,
None => {
let pool_meta = self.pool_meta.read().await;
(0..self.pools.len())
.map(|idx| pool_meta.is_suspended(idx))
.collect::<Vec<_>>()
}
};
let candidates = candidates.into_iter().flatten().collect();
let (winner_info, winner_idx) =
resolve_latest_object_info_candidates_with_pool_state(candidates, &suspended_pools, bucket, object, opts)?;
if let Some((prepared_idx, metadata)) = latest_prepared
&& prepared_idx == winner_idx
{
return Ok((metadata, winner_idx));
}
let pool = self.pools.get(winner_idx).ok_or(Error::ErasureReadQuorum)?;
#[cfg(test)]
pause_prepared_pool_read_before_refetch(object).await;
let metadata = pool
.prepare_get_object_reader_metadata(bucket, object, opts)
.await
.map_err(|err| to_object_err(err, vec![bucket, object]))?;
validate_prepared_pool_refetch_identity(&winner_info, metadata.object_info())?;
Ok((metadata, winner_idx))
}
pub(super) async fn delete_object_from_all_pools(
&self,
bucket: &str,
+31 -2
View File
@@ -218,7 +218,7 @@ fn same_user_defined_identity(left: &ObjectInfo, right: &ObjectInfo) -> bool {
/// excluded. The selected winner still carries the chosen pool's layout, while
/// the remaining read-visible fields must agree before the pool index can
/// provide a deterministic tie-break.
fn same_latest_object_info_identity(left: &ObjectInfo, right: &ObjectInfo) -> bool {
pub(super) fn same_latest_object_info_identity(left: &ObjectInfo, right: &ObjectInfo) -> bool {
let same_read_surface = left.bucket == right.bucket
&& left.name == right.name
&& left.is_dir == right.is_dir
@@ -277,6 +277,14 @@ fn same_latest_object_info_identity(left: &ObjectInfo, right: &ObjectInfo) -> bo
}
}
pub(super) fn validate_prepared_pool_refetch_identity(expected: &ObjectInfo, refetched: &ObjectInfo) -> Result<()> {
if same_latest_object_info_identity(expected, refetched) {
Ok(())
} else {
Err(Error::ErasureReadQuorum)
}
}
#[cfg(test)]
pub(super) fn resolve_latest_object_info_candidates(
candidates: Vec<LatestObjectInfoCandidate>,
@@ -328,7 +336,11 @@ pub(super) fn resolve_latest_object_info_candidates_with_pool_state(
return Err(Error::ErasureReadQuorum);
}
return Ok((winner_info.clone(), winner.idx));
let winner = latest_candidates.swap_remove(0);
let Some(winner_info) = winner.info else {
return Err(Error::ErasureReadQuorum);
};
return Ok((winner_info, winner.idx));
}
for candidate in candidates {
@@ -347,6 +359,23 @@ pub(super) fn resolve_latest_object_info_candidates_with_pool_state(
mod tests {
use super::*;
#[test]
fn prepared_pool_refetch_identity_fails_closed_on_generation_change() {
let expected = ObjectInfo {
mod_time: Some(OffsetDateTime::from_unix_timestamp(10).expect("test timestamp should be valid")),
version_id: Some(uuid::Uuid::from_u128(1)),
etag: Some("etag-a".to_string()),
..Default::default()
};
let mut refetched = expected.clone();
refetched.etag = Some("etag-b".to_string());
let error = validate_prepared_pool_refetch_identity(&expected, &refetched)
.expect_err("refetched metadata from a changed generation must fail closed");
assert_eq!(error, Error::ErasureReadQuorum);
}
#[test]
fn rebalance_delete_result_preserves_precondition_failed() {
let err = resolve_rebalance_delete_from_all_pools_result(Err(Error::PreconditionFailed), "bucket", "object")
@@ -0,0 +1,29 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use s3s::header::{
X_AMZ_OBJECT_LOCK_LEGAL_HOLD, X_AMZ_OBJECT_LOCK_MODE, X_AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE, X_AMZ_RESTORE,
X_AMZ_SERVER_SIDE_ENCRYPTION,
};
#[test]
fn persisted_metadata_keys_are_byte_stable() {
// These HTTP-header constants are also persisted xl.meta map keys. A drift can make an
// existing WORM retention appear absent or make a restored object's data directory reclaimable.
assert_eq!(X_AMZ_OBJECT_LOCK_LEGAL_HOLD.as_str(), "x-amz-object-lock-legal-hold");
assert_eq!(X_AMZ_OBJECT_LOCK_MODE.as_str(), "x-amz-object-lock-mode");
assert_eq!(X_AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE.as_str(), "x-amz-object-lock-retain-until-date");
assert_eq!(X_AMZ_RESTORE.as_str(), "x-amz-restore");
assert_eq!(X_AMZ_SERVER_SIDE_ENCRYPTION.as_str(), "x-amz-server-side-encryption");
}
+12 -2
View File
@@ -1537,7 +1537,7 @@ where
Ok(deleted_at)
}
pub async fn update_user_secret_key(&self, access_key: &str, secret_key: &str) -> Result<()> {
pub async fn update_user_secret_key(&self, access_key: &str, secret_key: &str) -> Result<(OffsetDateTime, AccountStatus)> {
if access_key.is_empty() || secret_key.is_empty() {
return Err(Error::InvalidArgument);
}
@@ -1552,7 +1552,16 @@ where
let mut cred = u.credentials.clone();
cred.secret_key = secret_key.to_string();
// Status is captured from the same credential snapshot the new secret
// is persisted with, so a caller replicating the rotation broadcasts
// exactly what was written rather than re-reading racily.
let status = if cred.is_valid() {
AccountStatus::Enabled
} else {
AccountStatus::Disabled
};
let u = UserIdentity::from(cred);
let updated_at = u.update_at.unwrap_or_else(OffsetDateTime::now_utc);
drop(cache);
drop(users);
@@ -1560,7 +1569,8 @@ where
.save_user_identity(access_key, UserType::Reg, u.clone(), None)
.await?;
self.update_user_with_claims(access_key, u)
self.update_user_with_claims(access_key, u)?;
Ok((updated_at, status))
}
/// Add SSH public key for a user (for SFTP authentication)
+8 -2
View File
@@ -960,7 +960,11 @@ impl<T: Store> IamSys<T> {
Ok(updated_at)
}
pub async fn set_user_secret_key(&self, access_key: &str, secret_key: &str) -> Result<()> {
pub async fn set_user_secret_key(
&self,
access_key: &str,
secret_key: &str,
) -> Result<(OffsetDateTime, rustfs_madmin::AccountStatus)> {
if !is_access_key_valid(access_key) {
return Err(IamError::InvalidAccessKeyLength);
}
@@ -969,7 +973,9 @@ impl<T: Store> IamSys<T> {
return Err(IamError::InvalidSecretKeyLength);
}
self.store.update_user_secret_key(access_key, secret_key).await
let (updated_at, status) = self.store.update_user_secret_key(access_key, secret_key).await?;
self.notify_for_user(access_key, false).await;
Ok((updated_at, status))
}
/// Add SSH public key for a user (for SFTP authentication)
+44 -3
View File
@@ -126,6 +126,27 @@ pub fn should_retry_delete_marker_purge(dobj: &DeletedObject) -> bool {
dobj.delete_marker_version_id.is_some()
}
/// True when the target denied a replicated delete because object-lock
/// retention or a legal hold protects that version on the replica (its
/// deletion gate answers `AccessDenied` with the lock reason, and a
/// replication request carries no governance bypass, rustfs#6850). Retrying
/// cannot succeed until the lock itself lapses, so callers treat this as a
/// policy denial rather than a transient fault.
///
/// The reason text is the RustFS deletion-gate wording; a MinIO/AWS peer
/// phrases its WORM denial differently and simply stays unclassified — the
/// caller then falls back to plain retry behavior, never a wrong state.
pub fn is_object_lock_denied_delete(code: Option<&str>, message: Option<&str>) -> bool {
if !matches!(code, Some("AccessDenied")) {
return false;
}
let Some(message) = message else {
return false;
};
let message = message.to_ascii_lowercase();
message.contains("retention") || message.contains("legal hold")
}
fn admitted_target_arns_from_replication_state(state: &ReplicationState) -> Vec<String> {
let mut target_arns = state.targets.keys().cloned().collect::<Vec<_>>();
target_arns.extend(state.purge_targets.keys().cloned());
@@ -237,9 +258,9 @@ mod tests {
use super::{
DeletedObjectReplicationInfo, delete_marker_purge_mrf_entry, delete_marker_purge_version_id,
delete_replication_creates_marker, is_retryable_delete_replication_head_error, is_version_delete_replication,
replicate_delete_outcome, resync_existing_delete_replication_info, should_retry_delete_marker_purge,
target_delete_version_id,
delete_replication_creates_marker, is_object_lock_denied_delete, is_retryable_delete_replication_head_error,
is_version_delete_replication, replicate_delete_outcome, resync_existing_delete_replication_info,
should_retry_delete_marker_purge, target_delete_version_id,
};
use crate::storage_api::DeletedObject;
use crate::{
@@ -615,4 +636,24 @@ mod tests {
corrupt.target_delete_marker_version_ids_corrupt = true;
assert_eq!(delete_marker_purge_version_id(Some(&corrupt), arn, source), None);
}
#[test]
fn object_lock_denied_delete_is_recognized_by_code_and_reason() {
// The peer's deletion gate answers AccessDenied with the lock reason.
assert!(is_object_lock_denied_delete(
Some("AccessDenied"),
Some("Object is under GOVERNANCE retention and cannot be deleted until 2026-09-01T00:00:00Z")
));
assert!(is_object_lock_denied_delete(
Some("AccessDenied"),
Some("Object has a legal hold and cannot be deleted. Remove the legal hold first.")
));
// A plain policy denial (misconfigured replicator) is not a lock denial.
assert!(!is_object_lock_denied_delete(Some("AccessDenied"), Some("Access Denied.")));
assert!(!is_object_lock_denied_delete(Some("AccessDenied"), None));
// Other errors mentioning retention must not match.
assert!(!is_object_lock_denied_delete(Some("InternalError"), Some("retention lookup failed")));
assert!(!is_object_lock_denied_delete(None, Some("legal hold")));
}
}
+5 -4
View File
@@ -41,9 +41,9 @@ pub use config::{
};
pub use delete::{
DeletedObjectReplicationInfo, delete_marker_purge_mrf_entry, delete_marker_purge_version_id,
delete_replication_creates_marker, is_retryable_delete_replication_head_error, is_version_delete_replication,
replicate_delete_outcome, resync_existing_delete_replication_info, should_retry_delete_marker_purge,
target_delete_version_id,
delete_replication_creates_marker, is_object_lock_denied_delete, is_retryable_delete_replication_head_error,
is_version_delete_replication, replicate_delete_outcome, resync_existing_delete_replication_info,
should_retry_delete_marker_purge, target_delete_version_id,
};
pub use filemeta::{
NULL_VERSION_ID, REPLICATE_EXISTING, REPLICATE_EXISTING_DELETE, REPLICATE_HEAL, REPLICATE_HEAL_DELETE, REPLICATE_INCOMING,
@@ -65,7 +65,8 @@ pub use multipart::{
pub use object::{
ReplicationSourceObject, ReplicationTargetObject, SsecPassthroughCapability, SsecPassthroughGate, content_matches_by_etag,
is_replication_target_offline_error, replication_action_for_target, replication_etags_match,
ssec_passthrough_evidence_present, ssec_passthrough_gate, target_is_newer_than_source_null_version, version_identity_drifted,
single_part_replica_etag_mismatch, ssec_passthrough_evidence_present, ssec_passthrough_gate,
target_is_newer_than_source_null_version, version_identity_drifted,
};
pub use operation::{
MustReplicateOptions, ReplicationDeleteScheduleInput, ReplicationDeleteSource, ReplicationDeleteStateSource,
+58 -2
View File
@@ -71,6 +71,32 @@ pub fn replication_etags_match(source: Option<&str>, target: Option<&str>) -> bo
source_etag.is_some() && source_etag == target_etag
}
fn is_plain_single_part_md5(etag: &str) -> bool {
etag.len() == 32 && etag.bytes().all(|b| b.is_ascii_hexdigit())
}
/// Whether the ETag the target returned for a single-part replica proves the
/// stored bytes differ from what the source sent — e.g. a target that does not
/// decode `aws-chunked` framing stores the frames verbatim and returns their
/// ETag. Only a plain single-part MD5 ETag on both sides is decidable; a
/// multipart or opaque (encrypted) ETag, or a withheld replica ETag, returns
/// `false` because no corruption can be concluded from it.
pub fn single_part_replica_etag_mismatch(source_etag: Option<&str>, replica_etag: Option<&str>) -> bool {
let Some(source) = source_etag.map(trim_etag) else {
return false;
};
if !is_plain_single_part_md5(&source) {
return false;
}
let Some(replica) = replica_etag.map(trim_etag) else {
return false;
};
if !is_plain_single_part_md5(&replica) {
return false;
}
!source.eq_ignore_ascii_case(&replica)
}
pub fn target_is_newer_than_source_null_version(
source: &ReplicationSourceObject<'_>,
target: &ReplicationTargetObject<'_>,
@@ -276,11 +302,41 @@ pub fn ssec_passthrough_evidence_present(sse_customer_algorithm: Option<&str>) -
#[cfg(test)]
mod tests {
const SOURCE_MD5: &str = "9a0364b9e99bb480dd25e1f0284c8555";
const FRAMED_MD5: &str = "0f343b0931126a20f133d67c2b018a3b";
#[test]
fn single_part_replica_mismatch_is_only_decided_on_plain_md5_pairs() {
// The #6853 shape: the target stored aws-chunked frames verbatim and
// returned the framed bytes' ETag.
assert!(single_part_replica_etag_mismatch(Some(SOURCE_MD5), Some(FRAMED_MD5)));
assert!(single_part_replica_etag_mismatch(
Some(&format!("\"{SOURCE_MD5}\"")),
Some(&format!("\"{FRAMED_MD5}\""))
));
// A faithful replica, quoted or not, passes; hex case must not matter
// (a target may return the same MD5 uppercased).
assert!(!single_part_replica_etag_mismatch(Some(SOURCE_MD5), Some(SOURCE_MD5)));
assert!(!single_part_replica_etag_mismatch(Some(&format!("\"{SOURCE_MD5}\"")), Some(SOURCE_MD5)));
assert!(!single_part_replica_etag_mismatch(
Some(SOURCE_MD5),
Some(&SOURCE_MD5.to_ascii_uppercase())
));
// Not decidable: multipart source, opaque replica ETag, or either side
// missing must never be reported as corruption.
assert!(!single_part_replica_etag_mismatch(Some(&format!("{SOURCE_MD5}-3")), Some(FRAMED_MD5)));
assert!(!single_part_replica_etag_mismatch(Some(SOURCE_MD5), Some(&format!("{FRAMED_MD5}-3"))));
assert!(!single_part_replica_etag_mismatch(Some(SOURCE_MD5), None));
assert!(!single_part_replica_etag_mismatch(None, Some(FRAMED_MD5)));
}
use super::{
ReplicationSourceObject, ReplicationTargetObject, SsecPassthroughCapability, SsecPassthroughGate,
content_matches_by_etag, is_replication_target_offline_error, replication_action_for_target, replication_etags_match,
ssec_passthrough_evidence_present, ssec_passthrough_gate, target_is_newer_than_source_null_version,
version_identity_drifted,
single_part_replica_etag_mismatch, ssec_passthrough_evidence_present, ssec_passthrough_gate,
target_is_newer_than_source_null_version, version_identity_drifted,
};
use crate::filemeta::{ReplicationAction, ReplicationType};
use crate::http::AMZ_OBJECT_LOCK_MODE;
+2
View File
@@ -23,10 +23,12 @@ use datafusion::{
use std::{error::Error as StdError, fmt::Display};
use thiserror::Error;
mod metrics;
pub mod object_store;
pub mod query;
pub mod server;
mod storage_api;
pub use metrics::{SelectInputMetrics, SelectInputMetricsSnapshot};
pub use storage_api::SelectObjectSnapshot;
#[cfg(test)]
+88
View File
@@ -0,0 +1,88 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use std::sync::atomic::{AtomicU64, Ordering};
#[derive(Clone, Copy, Debug, Default, Eq, PartialEq)]
pub struct SelectInputMetricsSnapshot {
pub bytes_scanned: u64,
pub bytes_processed: u64,
}
#[derive(Debug, Default)]
pub struct SelectInputMetrics {
uncompressed_bytes: AtomicU64,
}
impl SelectInputMetrics {
pub fn snapshot(&self) -> SelectInputMetricsSnapshot {
let uncompressed_bytes = self.uncompressed_bytes.load(Ordering::Relaxed);
SelectInputMetricsSnapshot {
bytes_scanned: uncompressed_bytes,
bytes_processed: uncompressed_bytes,
}
}
pub(crate) fn record_uncompressed(&self, bytes: usize) {
let increment = u64::try_from(bytes).unwrap_or(u64::MAX);
let _ = self
.uncompressed_bytes
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |current| Some(current.saturating_add(increment)));
}
/// Clears planner-only reads before query execution begins.
pub fn reset(&self) {
self.uncompressed_bytes.store(0, Ordering::Relaxed);
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn records_uncompressed_input_at_both_boundaries() {
let metrics = SelectInputMetrics::default();
metrics.record_uncompressed(7);
assert_eq!(
metrics.snapshot(),
SelectInputMetricsSnapshot {
bytes_scanned: 7,
bytes_processed: 7,
}
);
}
#[test]
fn counters_saturate_instead_of_wrapping() {
let metrics = SelectInputMetrics::default();
metrics.uncompressed_bytes.store(u64::MAX - 1, Ordering::Relaxed);
metrics.record_uncompressed(2);
assert_eq!(metrics.snapshot().bytes_scanned, u64::MAX);
assert_eq!(metrics.snapshot().bytes_processed, u64::MAX);
}
#[test]
fn reset_clears_schema_inference_bytes() {
let metrics = SelectInputMetrics::default();
metrics.record_uncompressed(9);
metrics.reset();
assert_eq!(metrics.snapshot(), SelectInputMetricsSnapshot::default());
}
}
File diff suppressed because it is too large Load Diff
+37
View File
@@ -13,6 +13,43 @@
// limitations under the License.
use datafusion::sql::sqlparser::ast::Statement;
use std::sync::Arc;
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum JsonPathSegment {
Key { name: String, quoted: bool },
Index(usize),
ArrayWildcard,
ObjectWildcard,
}
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub struct JsonSource {
path: Arc<[JsonPathSegment]>,
scalar_column: Option<String>,
}
impl JsonSource {
pub fn new(path: Vec<JsonPathSegment>, scalar_column: Option<String>) -> Self {
Self {
path: path.into(),
scalar_column,
}
}
#[cfg(test)]
pub(crate) fn from_path(path: Vec<JsonPathSegment>) -> Self {
Self::new(path, None)
}
pub fn path(&self) -> &[JsonPathSegment] {
&self.path
}
pub fn scalar_column(&self) -> Option<&str> {
self.scalar_column.as_deref()
}
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum ExtStatement {
+166
View File
@@ -25,6 +25,8 @@ use super::{
session::QueryAdmission,
};
pub type DispatchedQuery = (Query, Output);
#[async_trait]
pub trait QueryDispatcher: Send + Sync {
// fn create_query_id(&self) -> QueryId;
@@ -41,6 +43,18 @@ pub trait QueryDispatcher: Send + Sync {
self.execute_query(query).await
}
async fn dispatch_query(&self, query: &Query) -> QueryResult<DispatchedQuery> {
let execution_query = query.for_execution();
let output = self.execute_query(&execution_query).await?;
Ok((execution_query, output))
}
async fn dispatch_query_admitted(&self, query: &Query, admission: QueryAdmission) -> QueryResult<DispatchedQuery> {
let execution_query = query.for_execution();
let output = self.execute_query_admitted(&execution_query, admission).await?;
Ok((execution_query, output))
}
async fn build_logical_plan(&self, query_state_machine: Arc<QueryStateMachine>) -> QueryResult<Option<Plan>>;
async fn execute_logical_plan(&self, logical_plan: Plan, query_state_machine: Arc<QueryStateMachine>) -> QueryResult<Output>;
@@ -53,3 +67,155 @@ pub trait QueryDispatcher: Send + Sync {
// fn cancel_query(&self, id: &QueryId);
}
#[cfg(test)]
mod tests {
use super::*;
use crate::query::test_query;
use parking_lot::Mutex;
#[derive(Default)]
struct DefaultDispatchDispatcher {
executed_metrics: Mutex<Vec<Arc<crate::SelectInputMetrics>>>,
}
#[async_trait]
impl QueryDispatcher for DefaultDispatchDispatcher {
async fn execute_query(&self, query: &Query) -> QueryResult<Output> {
self.executed_metrics.lock().push(Arc::clone(query.input_metrics()));
Ok(Output::Nil(()))
}
async fn build_logical_plan(&self, _query_state_machine: Arc<QueryStateMachine>) -> QueryResult<Option<Plan>> {
unreachable!("default dispatch test does not plan queries")
}
async fn execute_logical_plan(
&self,
_logical_plan: Plan,
_query_state_machine: Arc<QueryStateMachine>,
) -> QueryResult<Output> {
unreachable!("default dispatch test does not execute plans")
}
async fn build_query_state_machine(&self, _query: Query) -> QueryResult<Arc<QueryStateMachine>> {
unreachable!("default dispatch test does not build state machines")
}
}
#[derive(Default)]
struct DistinctAdmittedDispatcher {
plain_metrics: Mutex<Vec<Arc<crate::SelectInputMetrics>>>,
admitted_metrics: Mutex<Vec<Arc<crate::SelectInputMetrics>>>,
fail_plain: bool,
fail_admitted: bool,
}
#[async_trait]
impl QueryDispatcher for DistinctAdmittedDispatcher {
async fn execute_query(&self, query: &Query) -> QueryResult<Output> {
self.plain_metrics.lock().push(Arc::clone(query.input_metrics()));
if self.fail_plain {
Err(crate::QueryError::Cancel)
} else {
Ok(Output::Nil(()))
}
}
async fn execute_query_admitted(&self, query: &Query, _admission: QueryAdmission) -> QueryResult<Output> {
self.admitted_metrics.lock().push(Arc::clone(query.input_metrics()));
if self.fail_admitted {
Err(crate::QueryError::Cancel)
} else {
Ok(Output::Nil(()))
}
}
async fn build_logical_plan(&self, _query_state_machine: Arc<QueryStateMachine>) -> QueryResult<Option<Plan>> {
unreachable!("dispatch routing test does not plan queries")
}
async fn execute_logical_plan(
&self,
_logical_plan: Plan,
_query_state_machine: Arc<QueryStateMachine>,
) -> QueryResult<Output> {
unreachable!("dispatch routing test does not execute plans")
}
async fn build_query_state_machine(&self, _query: Query) -> QueryResult<Arc<QueryStateMachine>> {
unreachable!("dispatch routing test does not build state machines")
}
}
#[tokio::test]
async fn plain_dispatch_propagates_override_errors() {
let dispatcher = DistinctAdmittedDispatcher {
fail_plain: true,
..Default::default()
};
let error = match dispatcher.dispatch_query(&test_query()).await {
Err(error) => error,
Ok(_) => panic!("plain override error should propagate"),
};
assert!(matches!(error, crate::QueryError::Cancel));
assert_eq!(dispatcher.plain_metrics.lock().len(), 1);
assert!(dispatcher.admitted_metrics.lock().is_empty());
}
#[tokio::test]
async fn default_dispatch_methods_use_distinct_execution_metrics() {
let dispatcher = DefaultDispatchDispatcher::default();
let query = test_query();
let (first, _) = dispatcher
.dispatch_query(&query)
.await
.expect("first dispatch should execute");
let (second, _) = dispatcher
.dispatch_query(&query)
.await
.expect("second dispatch should execute");
let (admitted, _) = dispatcher
.dispatch_query_admitted(&query, QueryAdmission::unmanaged())
.await
.expect("admitted dispatch should execute");
let executed_metrics = dispatcher.executed_metrics.lock();
assert!(!Arc::ptr_eq(first.input_metrics(), second.input_metrics()));
assert!(!Arc::ptr_eq(first.input_metrics(), admitted.input_metrics()));
assert!(Arc::ptr_eq(first.input_metrics(), &executed_metrics[0]));
assert!(Arc::ptr_eq(second.input_metrics(), &executed_metrics[1]));
assert!(Arc::ptr_eq(admitted.input_metrics(), &executed_metrics[2]));
}
#[tokio::test]
async fn admitted_dispatch_uses_the_admitted_override_and_propagates_errors() {
let dispatcher = DistinctAdmittedDispatcher::default();
let query = test_query();
let (dispatched, _) = dispatcher
.dispatch_query_admitted(&query, QueryAdmission::unmanaged())
.await
.expect("admitted dispatch should execute through its override");
assert!(dispatcher.plain_metrics.lock().is_empty());
{
let admitted_metrics = dispatcher.admitted_metrics.lock();
assert_eq!(admitted_metrics.len(), 1);
assert!(Arc::ptr_eq(dispatched.input_metrics(), &admitted_metrics[0]));
}
let failing = DistinctAdmittedDispatcher {
fail_admitted: true,
..Default::default()
};
let error = match failing.dispatch_query_admitted(&query, QueryAdmission::unmanaged()).await {
Err(error) => error,
Ok(_) => panic!("admitted override error should propagate"),
};
assert!(matches!(error, crate::QueryError::Cancel));
assert!(failing.plain_metrics.lock().is_empty());
assert_eq!(failing.admitted_metrics.lock().len(), 1);
}
}
@@ -29,6 +29,7 @@ use tracing::debug;
use crate::{QueryError, QueryResult};
use super::Query;
use super::ast::ExtStatement;
use super::logical_planner::Plan;
use super::session::{QueryExecutionTracker, SessionCtx};
@@ -172,6 +173,7 @@ pub struct QueryStateMachine {
pub session: SessionCtx,
pub query: Query,
prepared_statement: Option<ExtStatement>,
query_tracker: Option<QueryExecutionTracker>,
state: RwLock<QueryState>,
start: Instant,
@@ -196,6 +198,7 @@ impl QueryStateMachine {
Self {
session,
query,
prepared_statement: None,
query_tracker: None,
state: RwLock::new(QueryState::ACCEPTING),
start: Instant::now(),
@@ -211,6 +214,21 @@ impl QueryStateMachine {
Ok(state_machine)
}
pub fn begin_tracked_prepared(
query: Query,
session: SessionCtx,
query_tracker: QueryExecutionTracker,
prepared_statement: ExtStatement,
) -> QueryResult<Self> {
let mut state_machine = Self::begin_tracked(query, session, query_tracker)?;
state_machine.prepared_statement = Some(prepared_statement);
Ok(state_machine)
}
pub fn prepared_statement(&self) -> Option<&ExtStatement> {
self.prepared_statement.as_ref()
}
pub fn query_tracker(&self) -> Option<&QueryExecutionTracker> {
self.query_tracker.as_ref()
}
+46 -1
View File
@@ -15,7 +15,7 @@
use s3s::dto::SelectObjectContentInput;
use std::sync::Arc;
use crate::SelectObjectSnapshot;
use crate::{SelectInputMetrics, SelectObjectSnapshot};
pub mod analyzer;
pub mod ast;
@@ -40,6 +40,7 @@ pub struct Query {
context: Context,
content: String,
snapshot: Option<Arc<SelectObjectSnapshot>>,
input_metrics: Arc<SelectInputMetrics>,
}
impl Query {
@@ -49,6 +50,7 @@ impl Query {
context,
content,
snapshot: None,
input_metrics: Arc::new(SelectInputMetrics::default()),
}
}
@@ -58,6 +60,7 @@ impl Query {
context,
content,
snapshot: Some(snapshot),
input_metrics: Arc::new(SelectInputMetrics::default()),
}
}
@@ -72,4 +75,46 @@ impl Query {
pub fn snapshot(&self) -> Option<&Arc<SelectObjectSnapshot>> {
self.snapshot.as_ref()
}
pub fn input_metrics(&self) -> &Arc<SelectInputMetrics> {
&self.input_metrics
}
pub fn for_execution(&self) -> Self {
Self {
context: self.context.clone(),
content: self.content.clone(),
snapshot: self.snapshot.clone(),
input_metrics: Arc::new(SelectInputMetrics::default()),
}
}
}
#[cfg(test)]
fn test_query() -> Query {
use s3s::dto::{CSVInput, CSVOutput, ExpressionType, InputSerialization, OutputSerialization, SelectObjectContentRequest};
let input = SelectObjectContentInput {
bucket: "bucket".to_string(),
expected_bucket_owner: None,
key: "input.csv".to_string(),
sse_customer_algorithm: None,
sse_customer_key: None,
sse_customer_key_md5: None,
request: SelectObjectContentRequest {
expression: "SELECT * FROM S3Object".to_string(),
expression_type: ExpressionType::from_static(ExpressionType::SQL),
input_serialization: InputSerialization {
csv: Some(CSVInput::default()),
..Default::default()
},
output_serialization: OutputSerialization {
csv: Some(CSVOutput::default()),
..Default::default()
},
request_progress: None,
scan_range: None,
},
};
Query::new(Context { input: Arc::new(input) }, "SELECT * FROM S3Object".to_string())
}
+4
View File
@@ -34,6 +34,10 @@ impl Dialect for RustFsDialect {
fn supports_group_by_expr(&self) -> bool {
true
}
fn supports_partiql(&self) -> bool {
true
}
}
pub trait Parser {
+296 -19
View File
@@ -12,9 +12,11 @@
// See the License for the specific language governing permissions and
// limitations under the License.
use crate::SelectObjectSnapshot;
use crate::query::Context;
use crate::{QueryError, QueryResult, object_store::EcObjectStore};
use crate::query::{Context, Query, ast::JsonSource};
use crate::{
QueryError, QueryResult, SelectInputMetrics, SelectObjectSnapshot,
object_store::{EcObjectStore, is_json_document_input, legacy_json_source_from_input},
};
use datafusion::{
arrow::{
array::{Int32Array, StringArray},
@@ -314,8 +316,15 @@ impl SessionCtxFactory {
}
pub async fn create_session_ctx(&self, context: &Context) -> QueryResult<SessionCtx> {
self.create_session_ctx_inner(context, None, None, DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES)
.await
self.create_session_ctx_inner(
context,
None,
legacy_json_source_from_input(&context.input),
None,
None,
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES,
)
.await
}
pub async fn create_session_ctx_with_tracker_and_memory_limit(
@@ -324,8 +333,15 @@ impl SessionCtxFactory {
query_tracker: QueryExecutionTracker,
memory_limit_bytes: usize,
) -> QueryResult<SessionCtx> {
self.create_session_ctx_inner(context, None, Some(query_tracker), memory_limit_bytes)
.await
self.create_session_ctx_inner(
context,
None,
legacy_json_source_from_input(&context.input),
Some(query_tracker),
None,
memory_limit_bytes,
)
.await
}
pub async fn create_session_ctx_with_snapshot_and_tracker_and_memory_limit(
@@ -335,19 +351,63 @@ impl SessionCtxFactory {
query_tracker: QueryExecutionTracker,
memory_limit_bytes: usize,
) -> QueryResult<SessionCtx> {
self.create_session_ctx_inner(context, Some(snapshot), Some(query_tracker), memory_limit_bytes)
.await
self.create_session_ctx_inner(
context,
Some(snapshot),
legacy_json_source_from_input(&context.input),
Some(query_tracker),
None,
memory_limit_bytes,
)
.await
}
pub async fn create_session_ctx_for_query_with_source_and_tracker_and_memory_limit(
&self,
query: &Query,
source: JsonSource,
query_tracker: QueryExecutionTracker,
memory_limit_bytes: usize,
) -> QueryResult<SessionCtx> {
self.create_session_ctx_inner(
query.context(),
query.snapshot().cloned(),
source,
Some(query_tracker),
Some(Arc::clone(query.input_metrics())),
memory_limit_bytes,
)
.await
}
pub async fn create_session_ctx_for_query_with_tracker_and_memory_limit(
&self,
query: &Query,
query_tracker: QueryExecutionTracker,
memory_limit_bytes: usize,
) -> QueryResult<SessionCtx> {
self.create_session_ctx_inner(
query.context(),
query.snapshot().cloned(),
legacy_json_source_from_input(&query.context().input),
Some(query_tracker),
Some(Arc::clone(query.input_metrics())),
memory_limit_bytes,
)
.await
}
async fn create_session_ctx_inner(
&self,
context: &Context,
snapshot: Option<Arc<SelectObjectSnapshot>>,
source: JsonSource,
query_tracker: Option<QueryExecutionTracker>,
input_metrics: Option<Arc<SelectInputMetrics>>,
memory_limit_bytes: usize,
) -> QueryResult<SessionCtx> {
let df_session_ctx = self
.build_df_session_context(context, snapshot, query_tracker.clone(), memory_limit_bytes)
.build_df_session_context(context, snapshot, source, query_tracker.clone(), input_metrics, memory_limit_bytes)
.await?;
Ok(SessionCtx {
@@ -361,7 +421,9 @@ impl SessionCtxFactory {
&self,
context: &Context,
snapshot: Option<Arc<SelectObjectSnapshot>>,
source: JsonSource,
query_tracker: Option<QueryExecutionTracker>,
input_metrics: Option<Arc<SelectInputMetrics>>,
memory_limit_bytes: usize,
) -> QueryResult<SessionContext> {
let path = format!("s3://{}", context.input.bucket);
@@ -383,7 +445,14 @@ impl SessionCtxFactory {
.is_some_and(|delimiter| delimiter.len() == 2 && delimiter.as_bytes() != b"\r\n");
let scan_range_requires_single_file_scan =
context.input.request.scan_range.is_some() && context.input.request.input_serialization.parquet.is_none();
let config = if custom_two_byte_record_delimiter || scan_range_requires_single_file_scan {
let json_document_requires_single_file_scan = is_json_document_input(&context.input);
let metered_input_requires_single_file_scan =
input_metrics.is_some() && context.input.request.input_serialization.parquet.is_none();
let config = if custom_two_byte_record_delimiter
|| scan_range_requires_single_file_scan
|| json_document_requires_single_file_scan
|| metered_input_requires_single_file_scan
{
config.with_repartition_file_scans(false)
} else {
config
@@ -438,11 +507,23 @@ impl SessionCtxFactory {
df_session_state.with_object_store(&store_url, store).build()
} else {
let input_metrics = input_metrics.unwrap_or_else(|| Arc::new(SelectInputMetrics::default()));
let store: EcObjectStore = match query_tracker {
Some(query_tracker) => {
EcObjectStore::new_with_query_tracker(context.input.clone(), memory_pool, query_tracker, snapshot)
}
None => EcObjectStore::new_with_memory_pool(context.input.clone(), memory_pool, snapshot),
Some(query_tracker) => EcObjectStore::new_with_query_tracker_and_source(
context.input.clone(),
memory_pool,
query_tracker,
input_metrics,
snapshot,
source,
),
None => EcObjectStore::new_with_memory_pool_and_source(
context.input.clone(),
memory_pool,
input_metrics,
snapshot,
source,
),
}
.map_err(|err| QueryError::Datafusion {
source: Box::new(DataFusionError::External(Box::new(err))),
@@ -515,15 +596,15 @@ mod tests {
use crate::storage_api::object_store::ObjectIO as _;
use datafusion::{
datasource::{
file_format::csv::CsvFormat,
file_format::{csv::CsvFormat, json::JsonFormat},
listing::{ListingOptions, ListingTable, ListingTableConfig, ListingTableUrl},
},
execution::memory_pool::MemoryLimit,
};
use http::HeaderMap;
use s3s::dto::{
CSVInput, CSVOutput, ExpressionType, InputSerialization, JSONInput, OutputSerialization, ParquetInput, ScanRange,
SelectObjectContentInput, SelectObjectContentRequest,
CSVInput, CSVOutput, ExpressionType, InputSerialization, JSONInput, JSONType, OutputSerialization, ParquetInput,
ScanRange, SelectObjectContentInput, SelectObjectContentRequest,
};
use std::io::Write as _;
@@ -564,6 +645,103 @@ mod tests {
)
}
async fn test_query_tracker() -> QueryExecutionTracker {
let permit = Arc::new(tokio::sync::Semaphore::new(1))
.acquire_owned()
.await
.expect("query permit should be available");
QueryExecutionTracker::new(
&QueryExecutionOwner::new(),
Arc::new(permit),
Instant::now() + std::time::Duration::from_secs(300),
300,
)
}
async fn assert_legacy_json_column(bucket: &str, expression: &str, document: &[u8], column_name: &str, expected: &[&str]) {
const OBJECT: &str = "input.json";
let env = crate::storage_api::select_test_ecstore_env().await;
let mut context = test_context();
{
let input = Arc::make_mut(&mut context.input);
input.bucket = bucket.to_string();
input.key = OBJECT.to_string();
input.request.expression = expression.to_string();
input.request.input_serialization.csv = None;
input.request.input_serialization.json = Some(JSONInput {
type_: Some(JSONType::from_static(JSONType::DOCUMENT)),
});
}
env.make_bucket(bucket, false).await;
env.put_object_bytes(bucket, OBJECT, document.to_vec()).await;
let factory = SessionCtxFactory::new(false);
let lazy_session = factory
.create_session_ctx(&context)
.await
.expect("legacy lazy session should preserve the JSON source");
let tracked_lazy_session = factory
.create_session_ctx_with_tracker_and_memory_limit(
&context,
test_query_tracker().await,
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES,
)
.await
.expect("legacy tracked lazy session should preserve the JSON source");
let snapshot = prepare_test_snapshot(&context).await;
let snapshot_session = factory
.create_session_ctx_with_snapshot_and_tracker_and_memory_limit(
&context,
snapshot,
test_query_tracker().await,
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES,
)
.await
.expect("legacy snapshot session should preserve the JSON source");
for (kind, session) in [
("lazy", lazy_session),
("tracked lazy", tracked_lazy_session),
("tracked snapshot", snapshot_session),
] {
let table_path = ListingTableUrl::parse(format!("s3://{bucket}/{OBJECT}")).expect("parse JSON table URL");
let listing_options = ListingOptions::new(Arc::new(JsonFormat::default())).with_file_extension(".json");
let schema = listing_options
.infer_schema(session.inner(), &table_path)
.await
.expect("infer expanded JSON schema");
let table = ListingTable::try_new(
ListingTableConfig::new(table_path)
.with_listing_options(listing_options)
.with_schema(schema),
)
.expect("build expanded JSON table");
let query_context = SessionContext::new_with_state(session.inner().clone());
query_context
.register_table("legacy_input", Arc::new(table))
.expect("register expanded JSON table");
let batches = query_context
.sql(&format!("SELECT {column_name} FROM legacy_input"))
.await
.expect("plan expanded JSON query")
.collect()
.await
.expect("execute expanded JSON query");
let mut values = Vec::new();
for batch in batches {
let column = batch
.column(0)
.as_any()
.downcast_ref::<StringArray>()
.expect("expanded column should be Utf8");
for row in 0..batch.num_rows() {
values.push(column.value(row).to_string());
}
}
assert_eq!(values, expected, "{kind} legacy constructor");
}
}
#[test]
fn session_factory_fields_remain_source_compatible() {
let factory = SessionCtxFactory {
@@ -587,6 +765,32 @@ mod tests {
assert!(session.inner().config().options().optimizer.repartition_file_scans);
}
#[tokio::test]
async fn metered_csv_and_json_inputs_disable_file_repartitioning() {
let factory = SessionCtxFactory::new(true).with_target_partitions(3);
let csv_context = test_context();
let mut json_context = test_context();
let json_request = &mut Arc::make_mut(&mut json_context.input).request;
json_request.input_serialization.csv = None;
json_request.input_serialization.json = Some(JSONInput::default());
for context in [&csv_context, &json_context] {
let session = factory
.create_session_ctx_inner(
context,
None,
JsonSource::default(),
None,
Some(Arc::new(SelectInputMetrics::default())),
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES,
)
.await
.expect("metered session should be created");
assert!(!session.inner().config().options().optimizer.repartition_file_scans);
}
}
#[tokio::test]
async fn parquet_scan_range_keeps_file_repartitioning() {
let mut context = test_context();
@@ -627,6 +831,40 @@ mod tests {
assert!(!session.inner().config().options().optimizer.repartition_file_scans);
}
#[tokio::test]
async fn json_lines_without_scan_range_keeps_file_repartitioning() {
let mut context = test_context();
let request = &mut Arc::make_mut(&mut context.input).request;
request.input_serialization.csv = None;
request.input_serialization.json = Some(JSONInput::default());
let session = SessionCtxFactory::new(true)
.with_target_partitions(2)
.create_session_ctx(&context)
.await
.expect("JSON LINES session should be created");
assert!(session.inner().config().options().optimizer.repartition_file_scans);
}
#[tokio::test]
async fn json_document_disables_file_repartitioning() {
let mut context = test_context();
let request = &mut Arc::make_mut(&mut context.input).request;
request.input_serialization.csv = None;
request.input_serialization.json = Some(JSONInput {
type_: Some(JSONType::from_static(JSONType::DOCUMENT)),
});
let session = SessionCtxFactory::new(true)
.with_target_partitions(2)
.create_session_ctx(&context)
.await
.expect("JSON DOCUMENT session should be created");
assert!(!session.inner().config().options().optimizer.repartition_file_scans);
}
#[tokio::test]
async fn csv_scan_range_disables_file_repartitioning() {
let mut context = test_context();
@@ -702,7 +940,7 @@ mod tests {
async fn session_factory_applies_memory_limit() {
let factory = SessionCtxFactory::new(true);
let session = factory
.create_session_ctx_inner(&test_context(), None, None, 1024)
.create_session_ctx_inner(&test_context(), None, JsonSource::default(), None, None, 1024)
.await
.expect("session should be created with a bounded memory pool");
@@ -750,6 +988,45 @@ mod tests {
assert!(session.is_bound_to(&tracker));
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
#[serial_test::serial]
async fn legacy_session_factory_preserves_single_key_json_source() {
assert_legacy_json_column(
"s3select-legacy-session-json-source",
"SELECT e.name FROM S3Object.employees AS e",
br#"{"employees":[{"name":"Alice"}]}"#,
"name",
&["Alice"],
)
.await;
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
#[serial_test::serial]
async fn legacy_session_factory_preserves_implicit_root_alias() {
assert_legacy_json_column(
"s3select-legacy-session-root-alias",
"SELECT S3Object FROM S3Object",
br#"["one","two"]"#,
"s3object",
&["one", "two"],
)
.await;
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
#[serial_test::serial]
async fn legacy_session_factory_preserves_quoted_root_alias() {
assert_legacy_json_column(
"s3select-legacy-session-quoted-root-alias",
"SELECT \"V\" FROM S3Object AS \"V\"",
br#"["one","two"]"#,
"\"V\"",
&["one", "two"],
)
.await;
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
#[serial_test::serial]
async fn session_factory_propagates_query_guard_to_ec_store() {
+389 -50
View File
@@ -41,11 +41,11 @@ use rustfs_s3select_api::{
QueryError, QueryResult, SelectError,
query::{
Query,
ast::ExtStatement,
dispatcher::QueryDispatcher,
ast::{ExtStatement, JsonPathSegment, JsonSource},
dispatcher::{DispatchedQuery, QueryDispatcher},
execution::{Output, QueryStateMachine},
function::FuncMetaManagerRef,
logical_planner::{LogicalPlanner, Plan},
logical_planner::Plan,
parser::Parser,
session::{
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES, QueryAdmission, QueryExecutionOwner, QueryExecutionStatus,
@@ -53,7 +53,7 @@ use rustfs_s3select_api::{
},
},
};
use s3s::dto::{FileHeaderInfo, SelectObjectContentInput};
use s3s::dto::{FileHeaderInfo, JSONType, SelectObjectContentInput};
use std::sync::LazyLock;
use tokio::{
sync::Semaphore,
@@ -66,6 +66,7 @@ use crate::{
instance::{DEFAULT_MAX_CONCURRENT_QUERIES, DEFAULT_QUERY_TIMEOUT_SECS},
metadata::{ContextProviderExtension, MetadataProvider, TableHandleProviderRef, base_table::BaseTableProvider},
sql::logical::planner::DefaultLogicalPlanner,
sql::planner::prepare_s3_select_statement,
};
static IGNORE: LazyLock<FileHeaderInfo> = LazyLock::new(|| FileHeaderInfo::from_static(FileHeaderInfo::IGNORE));
@@ -120,7 +121,7 @@ impl Drop for QueryPhaseGuard<'_> {
#[async_trait]
impl QueryDispatcher for SimpleQueryDispatcher {
async fn execute_query(&self, query: &Query) -> QueryResult<Output> {
self.execute_query_inner(query, None).await
self.execute_query_inner(query, None).await.map(|(_, output)| output)
}
fn try_reserve_query(&self) -> QueryResult<QueryAdmission> {
@@ -133,6 +134,16 @@ impl QueryDispatcher for SimpleQueryDispatcher {
}
async fn execute_query_admitted(&self, query: &Query, admission: QueryAdmission) -> QueryResult<Output> {
self.execute_query_inner(query, Some(admission))
.await
.map(|(_, output)| output)
}
async fn dispatch_query(&self, query: &Query) -> QueryResult<DispatchedQuery> {
self.execute_query_inner(query, None).await
}
async fn dispatch_query_admitted(&self, query: &Query, admission: QueryAdmission) -> QueryResult<DispatchedQuery> {
self.execute_query_inner(query, Some(admission)).await
}
@@ -150,32 +161,26 @@ impl QueryDispatcher for SimpleQueryDispatcher {
let session = &query_state_machine.session;
let query = &query_state_machine.query;
let scheme_provider = self.build_scheme_provider(session).await?;
let logical_planner = DefaultLogicalPlanner::new(&scheme_provider);
let statements = self.parser.parse(query.content())?;
if statements.len() > 1 {
return Err(QueryError::MultiStatement {
num: statements.len(),
sql: query_state_machine.query.content().to_string(),
});
}
let stmt = match statements.front() {
Some(stmt) => stmt.clone(),
let stmt = match query_state_machine.prepared_statement() {
Some(statement) => statement.clone(),
None => {
return Err(QueryError::Parser {
source: ParserError::ParserError("empty SQL expression".to_string()),
});
let (statement, source) = self.prepare_query_statement(query.content())?;
if source_path_requires_expansion(source.path()) {
return Err(SelectError::DataSourcePathUnsupported.into());
}
statement
}
};
let scheme_provider = self.build_scheme_provider(session).await?;
let logical_planner = DefaultLogicalPlanner::new(&scheme_provider);
let logical_plan = self
.statement_to_logical_plan(stmt, &logical_planner, query_state_machine)
.statement_to_logical_plan(stmt, &logical_planner, Arc::clone(&query_state_machine))
.await?;
Ok(logical_plan)
})
.await?;
query_state_machine.query.input_metrics().reset();
if !query_tracker.mark_planned(&self.query_execution_owner) {
drop(logical_plan);
return Err(self.query_tracker_error(&query_tracker));
@@ -212,19 +217,21 @@ impl QueryDispatcher for SimpleQueryDispatcher {
}
async fn build_query_state_machine(&self, query: Query) -> QueryResult<Arc<QueryStateMachine>> {
self.build_query_state_machine_inner(query, None).await
self.build_query_state_machine_inner(query.for_execution(), None).await
}
}
impl SimpleQueryDispatcher {
async fn execute_query_inner(&self, query: &Query, admission: Option<QueryAdmission>) -> QueryResult<Output> {
let query_state_machine = self.build_query_state_machine_inner(query.clone(), admission).await?;
async fn execute_query_inner(&self, query: &Query, admission: Option<QueryAdmission>) -> QueryResult<DispatchedQuery> {
let query_state_machine = self.build_query_state_machine_inner(query.for_execution(), admission).await?;
let execution_query = query_state_machine.query.clone();
let logical_plan = self.build_logical_plan(Arc::clone(&query_state_machine)).await?;
let Some(logical_plan) = logical_plan else {
return Ok(Output::Nil(()));
return Ok((execution_query, Output::Nil(())));
};
self.execute_logical_plan(logical_plan, query_state_machine).await
let output = self.execute_logical_plan(logical_plan, query_state_machine).await?;
Ok((execution_query, output))
}
async fn build_query_state_machine_inner(
@@ -256,35 +263,52 @@ impl SimpleQueryDispatcher {
self.query_timeout.as_secs(),
);
let phase_guard = QueryPhaseGuard::new(&query_tracker, &self.query_execution_owner);
let session = if let Some(snapshot) = query.snapshot().cloned() {
self.run_with_query_deadline(
// Keep parser and analyzer errors in the planning phase. Successful
// preparation is cached here because the source path configures the
// object store before schema inference starts.
let (prepared_statement, source) = match self.prepare_query_statement(query.content()) {
Ok((statement, source)) => (Some(statement), source),
Err(_) => (None, JsonSource::default()),
};
let session = self
.run_with_query_deadline(
&query_tracker,
self.session_factory
.create_session_ctx_with_snapshot_and_tracker_and_memory_limit(
query.context(),
snapshot,
.create_session_ctx_for_query_with_source_and_tracker_and_memory_limit(
&query,
source,
query_tracker.clone(),
self.memory_limit_bytes,
),
)
.await?
} else {
self.run_with_query_deadline(
&query_tracker,
self.session_factory.create_session_ctx_with_tracker_and_memory_limit(
query.context(),
query_tracker.clone(),
self.memory_limit_bytes,
),
)
.await?
};
.await?;
if !query_tracker.mark_admitted(&self.query_execution_owner) {
drop(session);
return Err(self.query_tracker_error(&query_tracker));
}
phase_guard.disarm();
Ok(Arc::new(QueryStateMachine::begin_tracked(query, session, query_tracker)?))
let state_machine = match prepared_statement {
Some(statement) => QueryStateMachine::begin_tracked_prepared(query, session, query_tracker, statement)?,
None => QueryStateMachine::begin_tracked(query, session, query_tracker)?,
};
Ok(Arc::new(state_machine))
}
fn prepare_query_statement(&self, sql: &str) -> QueryResult<(ExtStatement, JsonSource)> {
let mut statements = self.parser.parse(sql)?;
if statements.len() > 1 {
return Err(QueryError::MultiStatement {
num: statements.len(),
sql: sql.to_string(),
});
}
let mut statement = statements.pop_front().ok_or_else(|| QueryError::Parser {
source: ParserError::ParserError("empty SQL expression".to_string()),
})?;
let ExtStatement::SqlStatement(sql_statement) = &mut statement;
let source = prepare_s3_select_statement(sql_statement)?;
validate_json_source_path_input(&self.input, source.path())?;
Ok((statement, source))
}
async fn run_with_query_deadline<T>(
&self,
@@ -364,7 +388,7 @@ impl SimpleQueryDispatcher {
// begin analyze
query_state_machine.begin_analyze();
let logical_plan = logical_planner
.create_logical_plan(stmt, &query_state_machine.session)
.prepared_statement_to_plan(stmt, &query_state_machine.session)
.await?;
query_state_machine.end_analyze();
@@ -502,6 +526,31 @@ impl SimpleQueryDispatcher {
}
}
fn validate_json_source_path_input(input: &SelectObjectContentInput, source_path: &[JsonPathSegment]) -> QueryResult<()> {
if source_path.is_empty() {
return Ok(());
}
let Some(json) = input.request.input_serialization.json.as_ref() else {
return Err(SelectError::DataSourcePathUnsupported.into());
};
if !source_path_requires_expansion(source_path)
|| json
.type_
.as_ref()
.is_some_and(|json_type| json_type.as_str() == JSONType::DOCUMENT)
{
return Ok(());
}
Err(SelectError::DataSourcePathUnsupported.into())
}
fn source_path_requires_expansion(source_path: &[JsonPathSegment]) -> bool {
!source_path
.strip_prefix(&[JsonPathSegment::ArrayWildcard])
.unwrap_or(source_path)
.is_empty()
}
pub struct TrackedRecordBatchStream {
state: Arc<TrackedRecordBatchState>,
schema: SchemaRef,
@@ -759,7 +808,10 @@ impl SimpleQueryDispatcherBuilder {
#[cfg(test)]
mod tests {
use super::{QueryPhaseGuard, SimpleQueryDispatcher, SimpleQueryDispatcherBuilder, TrackedRecordBatchStream};
use super::{
QueryPhaseGuard, SimpleQueryDispatcher, SimpleQueryDispatcherBuilder, TrackedRecordBatchStream,
validate_json_source_path_input,
};
use crate::{
execution::{
factory::{QueryExecutionFactoryRef, SqlQueryExecutionFactory},
@@ -788,6 +840,7 @@ mod tests {
QueryError, QueryResult, SelectError,
query::{
Context as QueryContext, Query,
ast::JsonPathSegment,
dispatcher::QueryDispatcher,
execution::{
Output, QueryExecution, QueryExecutionFactory, QueryExecutionRef, QueryStateMachine, QueryStateMachineRef,
@@ -1030,7 +1083,17 @@ mod tests {
query_execution_factory: QueryExecutionFactoryRef,
) -> (Arc<SimpleQueryDispatcher>, Arc<SelectObjectContentInput>) {
let input = Arc::new(test_input());
let dispatcher = SimpleQueryDispatcherBuilder::default()
let dispatcher = test_dispatcher_for_input(Arc::clone(&input), admission, query_timeout, query_execution_factory);
(dispatcher, input)
}
fn test_dispatcher_for_input(
input: Arc<SelectObjectContentInput>,
admission: Arc<Semaphore>,
query_timeout: Duration,
query_execution_factory: QueryExecutionFactoryRef,
) -> Arc<SimpleQueryDispatcher> {
SimpleQueryDispatcherBuilder::default()
.with_input(Arc::clone(&input))
.with_default_table_provider(Arc::new(BaseTableProvider::default()))
.with_session_factory(Arc::new(SessionCtxFactory::new(true)))
@@ -1040,8 +1103,7 @@ mod tests {
.with_query_admission(admission)
.with_query_timeout(query_timeout)
.build()
.expect("query dispatcher should build");
(dispatcher, input)
.expect("query dispatcher should build")
}
async fn snapshot_test_env() -> &'static TestECStoreEnv {
@@ -1170,6 +1232,68 @@ mod tests {
})
}
#[test]
fn nested_source_paths_require_json_document_input() {
let lines_input = json_snapshot_input();
let nested_path = [JsonPathSegment::Key {
name: "employees".to_string(),
quoted: false,
}];
assert!(matches!(
validate_json_source_path_input(&lines_input, &nested_path),
Err(ref error) if matches!(error.s3_select_policy_error(), Some(SelectError::DataSourcePathUnsupported))
));
assert!(validate_json_source_path_input(&lines_input, &[JsonPathSegment::ArrayWildcard]).is_ok());
let csv_input = test_input();
let parquet_input = parquet_snapshot_input();
for input in [&csv_input, parquet_input.as_ref()] {
assert!(matches!(
validate_json_source_path_input(input, &nested_path),
Err(ref error) if matches!(error.s3_select_policy_error(), Some(SelectError::DataSourcePathUnsupported))
));
}
let mut document_input = (*lines_input).clone();
document_input.request.input_serialization.json.as_mut().unwrap().type_ = Some(JSONType::from_static(JSONType::DOCUMENT));
assert!(validate_json_source_path_input(&document_input, &nested_path).is_ok());
}
#[tokio::test]
async fn normal_planning_pipeline_rejects_json_lines_source_expansion() {
let mut input = (*json_snapshot_input()).clone();
input.request.expression = "SELECT * FROM S3Object.employees".to_string();
let input = Arc::new(input);
let admission = Arc::new(Semaphore::new(1));
let optimizer = Arc::new(CascadeOptimizerBuilder::default().build());
let scheduler = Arc::new(LocalScheduler {});
let dispatcher = test_dispatcher_for_input(
Arc::clone(&input),
Arc::clone(&admission),
Duration::from_secs(300),
Arc::new(SqlQueryExecutionFactory::new(optimizer, scheduler)),
);
let query = Query::new(
QueryContext {
input: Arc::clone(&input),
},
input.request.expression.clone(),
);
let query_state_machine = dispatcher
.build_query_state_machine(query)
.await
.expect("validation errors should remain in the planning phase");
let result = dispatcher.build_logical_plan(query_state_machine).await;
assert!(matches!(
result,
Err(ref error) if matches!(error.s3_select_policy_error(), Some(SelectError::DataSourcePathUnsupported))
));
assert_eq!(admission.available_permits(), 1);
}
fn parquet_snapshot_input() -> Arc<SelectObjectContentInput> {
Arc::new(SelectObjectContentInput {
bucket: "s3select-parquet-snapshot-race".to_string(),
@@ -1529,6 +1653,140 @@ mod tests {
assert!(matches!(result, Err(QueryError::Cancel)));
}
#[tokio::test]
async fn unprepared_tracked_session_rejects_source_path_expansion() {
let mut input = test_input();
input.key = "test.json".to_string();
input.request.expression = "SELECT * FROM S3Object.employees".to_string();
input.request.input_serialization = InputSerialization {
json: Some(JSONInput {
type_: Some(JSONType::from_static(JSONType::DOCUMENT)),
}),
..Default::default()
};
input.request.output_serialization = OutputSerialization {
json: Some(JSONOutput::default()),
..Default::default()
};
let input = Arc::new(input);
let admission = Arc::new(Semaphore::new(1));
let optimizer = Arc::new(CascadeOptimizerBuilder::default().build());
let scheduler = Arc::new(LocalScheduler {});
let dispatcher = test_dispatcher_for_input(
Arc::clone(&input),
Arc::clone(&admission),
Duration::from_secs(300),
Arc::new(SqlQueryExecutionFactory::new(optimizer, scheduler)),
);
let query = Query::new(
QueryContext {
input: Arc::clone(&input),
},
input.request.expression.clone(),
);
let permit = Arc::clone(&admission).acquire_owned().await.expect("admission permit");
let tracker = QueryExecutionTracker::new(
&dispatcher.query_execution_owner,
Arc::new(permit),
Instant::now() + Duration::from_secs(300),
300,
);
let session = SessionCtxFactory::new(true)
.create_session_ctx_with_tracker_and_memory_limit(
query.context(),
tracker.clone(),
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES,
)
.await
.expect("test session");
assert!(tracker.mark_admitted(&dispatcher.query_execution_owner));
let state_machine =
Arc::new(QueryStateMachine::begin_tracked(query, session, tracker).expect("tracked state machine should be valid"));
let result = dispatcher.build_logical_plan(state_machine).await;
assert!(matches!(
result,
Err(ref error) if matches!(error.s3_select_policy_error(), Some(SelectError::DataSourcePathUnsupported))
));
assert_eq!(admission.available_permits(), 1);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn unprepared_tracked_session_preserves_root_scalar_bindings() {
for (bucket, expression) in [
("s3select-unprepared-root-scalar-alias", "SELECT V FROM S3Object AS V"),
("s3select-unprepared-root-wildcard", "SELECT _1 FROM S3Object[*]"),
] {
let mut input = test_input();
input.bucket = bucket.to_string();
input.key = "input.json".to_string();
input.request.expression = expression.to_string();
input.request.input_serialization = InputSerialization {
json: Some(JSONInput {
type_: Some(JSONType::from_static(JSONType::DOCUMENT)),
}),
..Default::default()
};
input.request.output_serialization = OutputSerialization {
json: Some(JSONOutput::default()),
..Default::default()
};
let input = Arc::new(input);
let env = snapshot_test_env().await;
env.make_bucket(&input.bucket, false).await;
env.put_object_bytes(&input.bucket, &input.key, br#"["one","two"]"#.to_vec())
.await;
let snapshot = env.prepare_select_object_snapshot(&input.bucket, &input.key).await;
let dispatcher = production_dispatcher(Arc::clone(&input));
let query = Query::new_with_snapshot(
QueryContext {
input: Arc::clone(&input),
},
input.request.expression.clone(),
Arc::clone(&snapshot),
);
let permit = Arc::clone(&dispatcher.query_admission)
.acquire_owned()
.await
.expect("query permit should be available");
let tracker = QueryExecutionTracker::new(
&dispatcher.query_execution_owner,
Arc::new(permit),
Instant::now() + Duration::from_secs(300),
300,
);
let session = SessionCtxFactory::new(false)
.create_session_ctx_with_snapshot_and_tracker_and_memory_limit(
query.context(),
snapshot,
tracker.clone(),
DEFAULT_S3SELECT_MEMORY_LIMIT_BYTES,
)
.await
.expect("legacy session should preserve the root scalar binding");
assert!(tracker.mark_admitted(&dispatcher.query_execution_owner));
let state_machine = Arc::new(
QueryStateMachine::begin_tracked(query, session, tracker).expect("tracked state machine should be valid"),
);
let logical_plan = dispatcher
.build_logical_plan(Arc::clone(&state_machine))
.await
.expect("unprepared query should plan")
.expect("SELECT should produce a logical plan");
let values = collect_utf8_output(
dispatcher
.execute_logical_plan(logical_plan, state_machine)
.await
.expect("unprepared query should execute"),
)
.await;
assert_eq!(values, ["one", "two"]);
}
}
#[tokio::test]
async fn staged_query_rejects_unbound_session() {
let admission = Arc::new(Semaphore::new(1));
@@ -1705,6 +1963,55 @@ mod tests {
assert_eq!(admission.available_permits(), 1);
}
#[tokio::test]
async fn reused_query_gets_execution_local_input_metrics() {
let admission = Arc::new(Semaphore::new(2));
let (dispatcher, input) = test_dispatcher(Arc::clone(&admission), Duration::from_secs(300));
let query = Query::new(QueryContext { input }, "SELECT * FROM S3Object".to_string());
let first = dispatcher
.build_query_state_machine(query.clone())
.await
.expect("first execution state");
let second = dispatcher
.build_query_state_machine(query)
.await
.expect("second execution state");
assert!(!Arc::ptr_eq(first.query.input_metrics(), second.query.input_metrics()));
drop((first, second));
assert_eq!(admission.available_permits(), 2);
}
#[tokio::test]
async fn dispatching_a_reused_query_returns_execution_local_input_metrics() {
let env = snapshot_test_env().await;
let mut input = test_input();
input.bucket = "s3select-reused-query-metrics".to_string();
input.key = "input.csv".to_string();
let input = Arc::new(input);
env.make_bucket(&input.bucket, false).await;
env.put_object_bytes(&input.bucket, &input.key, b"name\nAlice\n".to_vec())
.await;
let snapshot = env.prepare_select_object_snapshot(&input.bucket, &input.key).await;
let dispatcher = production_dispatcher(Arc::clone(&input));
let query = Query::new_with_snapshot(
QueryContext {
input: Arc::clone(&input),
},
input.request.expression.clone(),
snapshot,
);
let (first_query, first_output) = dispatcher.dispatch_query(&query).await.expect("first dispatch should start");
let (second_query, second_output) = dispatcher.dispatch_query(&query).await.expect("second dispatch should start");
assert!(!Arc::ptr_eq(first_query.input_metrics(), second_query.input_metrics()));
assert!(!Arc::ptr_eq(query.input_metrics(), first_query.input_metrics()));
assert!(!Arc::ptr_eq(query.input_metrics(), second_query.input_metrics()));
drop((first_output, second_output));
}
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
async fn concurrent_planning_claim_has_single_winner() {
let admission = Arc::new(Semaphore::new(1));
@@ -1905,6 +2212,38 @@ mod tests {
));
}
#[tokio::test]
async fn invalid_sql_precedes_malformed_json_snapshot_read() {
let mut input = json_snapshot_input();
let input_mut = Arc::make_mut(&mut input);
input_mut.bucket = "s3select-invalid-sql-precedence".to_string();
input_mut.key = "malformed.json".to_string();
input_mut.request.expression = "SELECT * FROM".to_string();
input_mut.request.input_serialization.json.as_mut().unwrap().type_ = Some(JSONType::from_static(JSONType::DOCUMENT));
let env = snapshot_test_env().await;
env.make_bucket(&input.bucket, false).await;
env.put_object_bytes(&input.bucket, &input.key, b"{bad".to_vec()).await;
let snapshot = env.prepare_select_object_snapshot(&input.bucket, &input.key).await;
let dispatcher = production_dispatcher(Arc::clone(&input));
let query = Query::new_with_snapshot(
QueryContext {
input: Arc::clone(&input),
},
input.request.expression.clone(),
snapshot,
);
let query_state_machine = dispatcher
.build_query_state_machine(query)
.await
.expect("invalid SQL should remain a planning-phase error");
assert!(matches!(
dispatcher.build_logical_plan(query_state_machine).await,
Err(QueryError::Parser { .. })
));
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn cancelled_execution_start_drops_future_before_releasing_admission() {
let admission = Arc::new(Semaphore::new(1));
+93 -6
View File
@@ -69,15 +69,15 @@ where
}
async fn execute(&self, query: &Query) -> QueryResult<QueryHandle> {
let result = self.query_dispatcher.execute_query(query).await?;
let (query, result) = self.query_dispatcher.dispatch_query(query).await?;
Ok(QueryHandle::new(query.clone(), result))
Ok(QueryHandle::new(query, result))
}
async fn execute_admitted(&self, query: &Query, admission: QueryAdmission) -> QueryResult<QueryHandle> {
let result = self.query_dispatcher.execute_query_admitted(query, admission).await?;
let (query, result) = self.query_dispatcher.dispatch_query_admitted(query, admission).await?;
Ok(QueryHandle::new(query.clone(), result))
Ok(QueryHandle::new(query, result))
}
async fn build_query_state_machine(&self, query: Query) -> QueryResult<QueryStateMachineRef> {
@@ -247,8 +247,19 @@ pub async fn make_rustfsms_with_components(
mod tests {
use std::sync::Arc;
use async_trait::async_trait;
use datafusion::{arrow::util::pretty, assert_batches_eq};
use rustfs_s3select_api::query::{Context, Query};
use parking_lot::Mutex;
use rustfs_s3select_api::{
QueryResult, SelectInputMetrics,
query::{
Context, Query,
dispatcher::QueryDispatcher,
execution::{Output, QueryStateMachine},
logical_planner::Plan,
},
server::dbms::DatabaseManagerSystem,
};
use s3s::dto::{
CSVInput, CSVOutput, ExpressionType, FieldDelimiter, FileHeaderInfo, InputSerialization, OutputSerialization,
RecordDelimiter, SelectObjectContentInput, SelectObjectContentRequest,
@@ -257,10 +268,66 @@ mod tests {
use crate::get_global_db;
use super::{
DEFAULT_MAX_CONCURRENT_QUERIES, DEFAULT_MEMORY_LIMIT_BYTES, DEFAULT_QUERY_TIMEOUT_SECS, MAX_QUERY_TIMEOUT_SECS,
DEFAULT_MAX_CONCURRENT_QUERIES, DEFAULT_MEMORY_LIMIT_BYTES, DEFAULT_QUERY_TIMEOUT_SECS, MAX_QUERY_TIMEOUT_SECS, RustFSms,
S3SelectRuntimeConfig, bounded_u64_from_env_value, bounded_usize_from_env_value, target_partitions_from_env_value,
};
#[derive(Default)]
struct FreshMetricsDispatcher {
executed_metrics: Mutex<Vec<Arc<SelectInputMetrics>>>,
}
#[async_trait]
impl QueryDispatcher for FreshMetricsDispatcher {
async fn execute_query(&self, query: &Query) -> QueryResult<Output> {
self.executed_metrics.lock().push(Arc::clone(query.input_metrics()));
Ok(Output::Nil(()))
}
async fn build_logical_plan(&self, _query_state_machine: Arc<QueryStateMachine>) -> QueryResult<Option<Plan>> {
unreachable!("fresh metrics test does not plan queries")
}
async fn execute_logical_plan(
&self,
_logical_plan: Plan,
_query_state_machine: Arc<QueryStateMachine>,
) -> QueryResult<Output> {
unreachable!("fresh metrics test does not execute plans")
}
async fn build_query_state_machine(&self, _query: Query) -> QueryResult<Arc<QueryStateMachine>> {
unreachable!("fresh metrics test does not build state machines")
}
}
fn metrics_test_query() -> Query {
let expression = "SELECT * FROM S3Object";
let input = SelectObjectContentInput {
bucket: "bucket".to_string(),
expected_bucket_owner: None,
key: "input.csv".to_string(),
sse_customer_algorithm: None,
sse_customer_key: None,
sse_customer_key_md5: None,
request: SelectObjectContentRequest {
expression: expression.to_string(),
expression_type: ExpressionType::from_static(ExpressionType::SQL),
input_serialization: InputSerialization {
csv: Some(CSVInput::default()),
..Default::default()
},
output_serialization: OutputSerialization {
csv: Some(CSVOutput::default()),
..Default::default()
},
request_progress: None,
scan_range: None,
},
};
Query::new(Context { input: Arc::new(input) }, expression.to_string())
}
#[test]
fn parses_target_partitions_from_env_value() {
assert_eq!(target_partitions_from_env_value(Some("4")), 4);
@@ -291,6 +358,26 @@ mod tests {
assert_eq!(bounded_u64_from_env_value(None, 300, MAX_QUERY_TIMEOUT_SECS), 300);
}
#[tokio::test]
async fn repeated_execute_returns_the_fresh_dispatched_query_metrics() {
let dispatcher = Arc::new(FreshMetricsDispatcher::default());
let db = RustFSms {
query_dispatcher: Arc::clone(&dispatcher),
};
let query = metrics_test_query();
let first = db.execute(&query).await.expect("first execution should succeed");
let second = db.execute(&query).await.expect("second execution should succeed");
let executed_metrics = dispatcher.executed_metrics.lock();
assert_eq!(executed_metrics.len(), 2);
assert!(Arc::ptr_eq(first.query().input_metrics(), &executed_metrics[0]));
assert!(Arc::ptr_eq(second.query().input_metrics(), &executed_metrics[1]));
assert!(!Arc::ptr_eq(first.query().input_metrics(), second.query().input_metrics()));
assert!(!Arc::ptr_eq(first.query().input_metrics(), query.input_metrics()));
assert!(!Arc::ptr_eq(second.query().input_metrics(), query.input_metrics()));
}
#[tokio::test]
#[ignore = "requires a live RustFS store with a pre-seeded test object (bucket 'dandan')"]
async fn test_simple_sql() {
+7
View File
@@ -184,6 +184,13 @@ mod tests {
assert!(dialect.supports_group_by_expr(), "RustFsDialect should support GROUP BY expressions");
}
#[test]
fn test_supports_partiql_paths() {
let dialect = RustFsDialect;
assert!(dialect.supports_partiql(), "RustFsDialect should support JSON source paths");
}
#[test]
fn test_identifier_validation_comprehensive() {
let dialect = RustFsDialect;
+55 -1
View File
@@ -16,6 +16,7 @@ use std::{collections::VecDeque, fmt::Display};
use datafusion::sql::sqlparser::{
dialect::Dialect,
keywords::{Keyword, RESERVED_FOR_TABLE_ALIAS},
parser::{Parser, ParserError},
tokenizer::{Token, Tokenizer},
};
@@ -53,7 +54,8 @@ impl<'a> ExtParser<'a> {
/// Parse the specified tokens with dialect
fn new_with_dialect(sql: &str, dialect: &'a dyn Dialect) -> Result<Self> {
let mut tokenizer = Tokenizer::new(dialect, sql);
let tokens = tokenizer.tokenize()?;
let mut tokens = tokenizer.tokenize()?;
rewrite_source_object_wildcards(&mut tokens);
Ok(ExtParser {
parser: Parser::new(dialect).with_tokens(tokens),
})
@@ -104,6 +106,41 @@ impl<'a> ExtParser<'a> {
}
}
fn rewrite_source_object_wildcards(tokens: &mut [Token]) {
let mut paren_depth = 0usize;
let mut in_from = false;
let mut index = 0usize;
while index < tokens.len() {
match &tokens[index] {
Token::Word(word) if paren_depth == 0 && word.keyword == Keyword::FROM => {
in_from = true;
}
Token::Word(word) if in_from && paren_depth == 0 && ends_from_source(word.keyword) => {
in_from = false;
}
Token::SemiColon if paren_depth == 0 => in_from = false,
Token::LParen => paren_depth = paren_depth.saturating_add(1),
Token::RParen => paren_depth = paren_depth.saturating_sub(1),
Token::Period if in_from => {
if let Some(next) = tokens[index + 1..]
.iter_mut()
.find(|token| !matches!(token, Token::Whitespace(_)))
&& matches!(next, Token::Mul)
{
*next = Token::make_word("*", None);
}
}
_ => {}
}
index += 1;
}
}
fn ends_from_source(keyword: Keyword) -> bool {
RESERVED_FOR_TABLE_ALIAS.contains(&keyword) || matches!(keyword, Keyword::PREWHERE | Keyword::SETTINGS | Keyword::FORMAT)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -172,6 +209,23 @@ mod tests {
}
}
#[test]
fn parses_source_object_wildcard_without_rewriting_projection_wildcard() {
let mut statements = ExtParser::parse_sql("SELECT e.* FROM S3Object[*].* AS e").expect("query should parse");
let ExtStatement::SqlStatement(statement) = statements.pop_front().expect("one statement");
assert_eq!(statement.to_string(), "SELECT e.* FROM S3Object[*].* AS e");
}
#[test]
fn from_tokens_in_literals_and_comments_do_not_change_wildcard_scope() {
let sql = "SELECT 'FROM x.*' AS marker /* FROM y.* */ FROM S3Object.*";
let mut statements = ExtParser::parse_sql(sql).expect("query should parse");
let ExtStatement::SqlStatement(statement) = statements.pop_front().expect("one statement");
assert_eq!(statement.to_string(), "SELECT 'FROM x.*' AS marker FROM S3Object.*");
}
#[test]
fn test_default_parser_multiple_statements() {
let parser = DefaultParser::default();
+361 -33
View File
@@ -12,20 +12,21 @@
// See the License for the specific language governing permissions and
// limitations under the License.
use std::ops::ControlFlow;
use std::{convert::Infallible, ops::ControlFlow};
use async_recursion::async_recursion;
use async_trait::async_trait;
use datafusion::sql::{
planner::SqlToRel,
planner::{IdentNormalizer, SqlToRel},
sqlparser::ast::{
GroupByExpr, ObjectNamePart, OrderByKind, Query, Select, SelectFlavor, SetExpr, Statement, TableFactor, Visit, Visitor,
AccessExpr, Expr, GroupByExpr, Ident, JsonPath, JsonPathElem, ObjectNamePart, OrderByKind, Query, Select, SelectFlavor,
SetExpr, Statement, Subscript, TableAlias, TableFactor, Value, Visit, VisitMut, Visitor, VisitorMut,
},
};
use rustfs_s3select_api::{
QueryError, QueryResult, SelectError,
query::{
ast::ExtStatement,
ast::{ExtStatement, JsonPathSegment, JsonSource},
logical_planner::{LogicalPlanner, Plan, QueryPlan},
session::SessionCtx,
},
@@ -64,21 +65,24 @@ impl<'a, S: ContextProviderExtension + Send + Sync + 'a> SqlPlanner<'a, S> {
}
}
async fn df_sql_to_plan(&self, stmt: Statement, _session: &SessionCtx) -> QueryResult<Plan> {
match stmt {
Statement::Query(_) => {
validate_s3_select_statement(&stmt)?;
let df_plan = self.df_planner.sql_statement_to_plan(stmt).map_err(classify_planner_error)?;
let plan = Plan::Query(QueryPlan {
df_plan,
is_tag_scan: false,
});
Ok(plan)
}
_ => Err(unsupported_structure("only SELECT queries are supported")),
pub(crate) async fn prepared_statement_to_plan(&self, statement: ExtStatement, session: &SessionCtx) -> QueryResult<Plan> {
match statement {
ExtStatement::SqlStatement(stmt) => self.df_prepared_sql_to_plan(*stmt, session).await,
}
}
async fn df_sql_to_plan(&self, mut stmt: Statement, session: &SessionCtx) -> QueryResult<Plan> {
prepare_s3_select_statement(&mut stmt)?;
self.df_prepared_sql_to_plan(stmt, session).await
}
async fn df_prepared_sql_to_plan(&self, stmt: Statement, _session: &SessionCtx) -> QueryResult<Plan> {
let df_plan = self.df_planner.sql_statement_to_plan(stmt).map_err(classify_planner_error)?;
Ok(Plan::Query(QueryPlan {
df_plan,
is_tag_scan: false,
}))
}
}
fn classify_planner_error(error: datafusion::common::DataFusionError) -> QueryError {
@@ -95,11 +99,10 @@ fn classify_planner_error(error: datafusion::common::DataFusionError) -> QueryEr
error.into()
}
fn validate_s3_select_statement(statement: &Statement) -> QueryResult<()> {
pub(crate) fn prepare_s3_select_statement(statement: &mut Statement) -> QueryResult<JsonSource> {
let Statement::Query(query) = statement else {
return Err(unsupported_structure("only SELECT queries are supported"));
};
if query.with.is_some()
|| query.order_by.as_ref().is_some_and(|order_by| {
order_by.interpolate.is_some()
@@ -137,17 +140,81 @@ fn validate_s3_select_statement(statement: &Statement) -> QueryResult<()> {
}
let mut detector = SubqueryDetector { visited_root: false };
if query.visit(&mut detector).is_break() {
if Visit::visit(&*query, &mut detector).is_break() {
return Err(unsupported_structure("subqueries are not supported"));
}
let SetExpr::Select(select) = query.body.as_ref() else {
return Err(unsupported_structure("set operations and nested queries are not supported"));
let source = {
let SetExpr::Select(select) = query.body.as_mut() else {
return Err(unsupported_structure("set operations and nested queries are not supported"));
};
prepare_select(select)?
};
validate_select(select)
let mut normalizer = PartiQlSubscriptNormalizer;
let _ = VisitMut::visit(query, &mut normalizer);
Ok(source)
}
fn validate_select(select: &Select) -> QueryResult<()> {
struct PartiQlSubscriptNormalizer;
impl VisitorMut for PartiQlSubscriptNormalizer {
type Break = Infallible;
fn post_visit_expr(&mut self, expr: &mut Expr) -> ControlFlow<Self::Break> {
let Expr::JsonAccess { value, path } = expr else {
return ControlFlow::Continue(());
};
if !matches!(path.path.first(), Some(JsonPathElem::Bracket { .. }))
|| path
.path
.iter()
.any(|element| matches!(element, JsonPathElem::ColonBracket { .. }))
{
return ControlFlow::Continue(());
}
let mut appended_access = Vec::with_capacity(path.path.len());
for element in std::mem::take(&mut path.path) {
match element {
JsonPathElem::Dot { key, quoted } => {
let identifier = if quoted {
Ident::with_quote('"', key)
} else {
Ident::new(key)
};
appended_access.push(AccessExpr::Dot(Expr::Identifier(identifier)));
}
JsonPathElem::Bracket { key } => {
appended_access.push(AccessExpr::Subscript(Subscript::Index { index: key }));
}
JsonPathElem::ColonBracket { key } => {
appended_access.push(AccessExpr::Subscript(Subscript::Index { index: key }));
}
}
}
let value = std::mem::replace(value, Box::new(Expr::Identifier(Ident::new(""))));
let (root, mut access_chain) = match *value {
Expr::CompoundFieldAccess { root, access_chain } => (root, access_chain),
root => (Box::new(root), Vec::new()),
};
access_chain.extend(appended_access);
*expr = Expr::CompoundFieldAccess { root, access_chain };
ControlFlow::Continue(())
}
}
fn implicit_source_alias(source_path: &[JsonPathSegment]) -> Ident {
match source_path.last() {
Some(JsonPathSegment::Key { name, quoted: true }) => Ident::with_quote('"', name),
Some(JsonPathSegment::Key { name, quoted: false }) => Ident::new(name),
Some(JsonPathSegment::Index(_) | JsonPathSegment::ArrayWildcard | JsonPathSegment::ObjectWildcard) | None => {
Ident::new("_1")
}
}
}
fn prepare_select(select: &mut Select) -> QueryResult<JsonSource> {
if !select.optimizer_hints.is_empty()
|| select.distinct.is_some()
|| select.select_modifiers.is_some()
@@ -170,7 +237,7 @@ fn validate_select(select: &Select) -> QueryResult<()> {
return Err(unsupported_structure("the SELECT contains an unsupported clause"));
}
let [table] = select.from.as_slice() else {
let [table] = select.from.as_mut_slice() else {
return Err(unsupported_structure("exactly one S3Object source is required"));
};
if !table.joins.is_empty() {
@@ -186,8 +253,8 @@ fn validate_select(select: &Select) -> QueryResult<()> {
partitions,
sample,
index_hints,
..
} = &table.relation
json_path,
} = &mut table.relation
else {
return Err(unsupported_structure("subqueries and table functions are not supported"));
};
@@ -202,9 +269,7 @@ fn validate_select(select: &Select) -> QueryResult<()> {
{
return Err(unsupported_structure("the S3Object source contains unsupported modifiers"));
}
let ([ObjectNamePart::Identifier(table_name)] | [ObjectNamePart::Identifier(table_name), ObjectNamePart::Identifier(_)]) =
name.0.as_slice()
else {
let Some(ObjectNamePart::Identifier(table_name)) = name.0.first() else {
return Err(SelectError::DataSourcePathUnsupported.into());
};
let is_s3_object = if table_name.quote_style.is_some() {
@@ -216,6 +281,72 @@ fn validate_select(select: &Select) -> QueryResult<()> {
return Err(SelectError::DataSourcePathUnsupported.into());
}
let mut source_path = Vec::new();
for part in &name.0[1..] {
let ObjectNamePart::Identifier(identifier) = part else {
return Err(SelectError::DataSourcePathUnsupported.into());
};
if identifier.quote_style.is_none() && identifier.value == "*" {
source_path.push(JsonPathSegment::ObjectWildcard);
} else {
source_path.push(JsonPathSegment::Key {
name: identifier.value.clone(),
quoted: identifier.quote_style.is_some(),
});
}
}
if let Some(json_path) = json_path.as_ref() {
append_json_path_segments(&mut source_path, json_path)?;
}
if alias.is_none() && !source_path.is_empty() {
*alias = Some(TableAlias {
explicit: true,
name: implicit_source_alias(&source_path),
columns: Vec::new(),
at: None,
});
}
let scalar_column = alias
.as_ref()
.map(|alias| IdentNormalizer::default().normalize(alias.name.clone()))
.or_else(|| {
source_path
.is_empty()
.then(|| IdentNormalizer::default().normalize(table_name.clone()))
});
name.0.truncate(1);
*json_path = None;
Ok(JsonSource::new(source_path, scalar_column))
}
fn append_json_path_segments(source_path: &mut Vec<JsonPathSegment>, json_path: &JsonPath) -> QueryResult<()> {
for element in &json_path.path {
let segment = match element {
JsonPathElem::Dot { key, quoted } if key == "*" && !quoted => JsonPathSegment::ObjectWildcard,
JsonPathElem::Dot { key, quoted } => JsonPathSegment::Key {
name: key.clone(),
quoted: *quoted,
},
JsonPathElem::Bracket { key: Expr::Wildcard(_) } => JsonPathSegment::ArrayWildcard,
JsonPathElem::Bracket { key: Expr::Value(value) } => match &value.value {
Value::Number(number, false) => JsonPathSegment::Index(
number
.to_string()
.parse()
.map_err(|_| QueryError::from(SelectError::DataSourcePathUnsupported))?,
),
Value::SingleQuotedString(key) => JsonPathSegment::Key {
name: key.clone(),
quoted: true,
},
_ => return Err(SelectError::DataSourcePathUnsupported.into()),
},
JsonPathElem::Bracket { .. } | JsonPathElem::ColonBracket { .. } => {
return Err(SelectError::DataSourcePathUnsupported.into());
}
};
source_path.push(segment);
}
Ok(())
}
@@ -245,10 +376,14 @@ impl Visitor for SubqueryDetector {
#[cfg(test)]
mod tests {
use super::validate_s3_select_statement;
use super::prepare_s3_select_statement;
use crate::sql::parser::ExtParser;
use datafusion::sql::sqlparser::ast::Statement;
use rustfs_s3select_api::{SelectError, query::ast::ExtStatement};
use datafusion::sql::sqlparser::ast::{AccessExpr, Expr, Statement, Visit, Visitor};
use rustfs_s3select_api::{
QueryResult, SelectError,
query::ast::{ExtStatement, JsonPathSegment, JsonSource},
};
use std::ops::ControlFlow;
fn parse_statement(sql: &str) -> Statement {
let mut statements = ExtParser::parse_sql(sql).expect("SQL should parse");
@@ -256,6 +391,10 @@ mod tests {
*statement
}
fn validate_s3_select_statement(statement: &Statement) -> QueryResult<JsonSource> {
prepare_s3_select_statement(&mut statement.clone())
}
#[test]
fn accepts_s3_select_query_shape() {
let statement = parse_statement("SELECT s.id FROM S3Object AS s WHERE s.id = '1' LIMIT 10");
@@ -270,6 +409,195 @@ mod tests {
assert!(validate_s3_select_statement(&statement).is_ok());
}
#[test]
fn prepares_nested_json_source_path_and_normalizes_table() {
let mut statement = parse_statement("SELECT e.name FROM S3Object[*].employees[*] AS e");
let source = prepare_s3_select_statement(&mut statement).expect("JSON source path should be supported");
assert_eq!(
source.path(),
&[
JsonPathSegment::ArrayWildcard,
JsonPathSegment::Key {
name: "employees".to_string(),
quoted: false,
},
JsonPathSegment::ArrayWildcard,
]
);
assert_eq!(statement.to_string(), "SELECT e.name FROM S3Object AS e");
}
#[test]
fn partiql_source_support_preserves_projection_and_filter_subscripts() {
let mut statement = parse_statement("SELECT s.tags[1] FROM S3Object AS s WHERE s.values[0] = 1");
prepare_s3_select_statement(&mut statement).expect("array expressions should remain supported");
let mut counter = FieldAccessCounter::default();
let _ = Visit::visit(&statement, &mut counter);
assert_eq!(counter.json_accesses, 0);
assert_eq!(counter.subscripts, 2);
}
#[derive(Default)]
struct FieldAccessCounter {
json_accesses: usize,
subscripts: usize,
}
impl Visitor for FieldAccessCounter {
type Break = ();
fn pre_visit_expr(&mut self, expr: &Expr) -> ControlFlow<Self::Break> {
match expr {
Expr::JsonAccess { .. } => self.json_accesses += 1,
Expr::CompoundFieldAccess { access_chain, .. } => {
self.subscripts += access_chain
.iter()
.filter(|access| matches!(access, AccessExpr::Subscript(_)))
.count();
}
_ => {}
}
ControlFlow::Continue(())
}
}
#[test]
fn prepares_array_index_and_object_wildcard_paths() {
let mut index_statement = parse_statement("SELECT * FROM S3Object[0]");
let mut wildcard_statement = parse_statement("SELECT * FROM S3Object[*].*");
assert_eq!(
prepare_s3_select_statement(&mut index_statement)
.expect("array index should be supported")
.path(),
&[JsonPathSegment::Index(0)]
);
assert_eq!(
prepare_s3_select_statement(&mut wildcard_statement)
.expect("object wildcard should be supported")
.path(),
&[JsonPathSegment::ArrayWildcard, JsonPathSegment::ObjectWildcard]
);
}
#[test]
fn quoted_star_remains_an_object_key() {
let mut statement = parse_statement("SELECT * FROM S3Object.\"*\"");
let source = prepare_s3_select_statement(&mut statement).expect("quoted key should be supported");
assert_eq!(
source.path(),
&[JsonPathSegment::Key {
name: "*".to_string(),
quoted: true,
}]
);
}
#[test]
fn preserves_quoted_keys_and_adds_implicit_source_aliases() {
let mut key_statement = parse_statement("SELECT employee.name FROM S3Object[*].department.employee");
let mut wildcard_statement = parse_statement("SELECT _1.name FROM S3Object[*].employees[*]");
prepare_s3_select_statement(&mut key_statement).expect("named source path should be supported");
prepare_s3_select_statement(&mut wildcard_statement).expect("wildcard source path should be supported");
assert_eq!(key_statement.to_string(), "SELECT employee.name FROM S3Object AS employee");
assert_eq!(wildcard_statement.to_string(), "SELECT _1.name FROM S3Object AS _1");
}
#[test]
fn root_scalar_aliases_are_preserved_and_unquoted_aliases_are_normalized() {
let mut implicit = parse_statement("SELECT S3Object FROM S3Object");
let mut unquoted = parse_statement("SELECT V FROM S3Object AS V");
let mut quoted = parse_statement("SELECT \"V\" FROM S3Object AS \"V\"");
let implicit_source = prepare_s3_select_statement(&mut implicit).expect("implicit root alias should be supported");
let unquoted_source = prepare_s3_select_statement(&mut unquoted).expect("unquoted root alias should be supported");
let quoted_source = prepare_s3_select_statement(&mut quoted).expect("quoted root alias should be supported");
assert!(implicit_source.path().is_empty());
assert_eq!(implicit_source.scalar_column(), Some("s3object"));
assert!(unquoted_source.path().is_empty());
assert_eq!(unquoted_source.scalar_column(), Some("v"));
assert!(quoted_source.path().is_empty());
assert_eq!(quoted_source.scalar_column(), Some("V"));
}
#[test]
fn unquoted_terminal_scalar_alias_uses_datafusion_identifier_case() {
let mut statement = parse_statement("SELECT NAME FROM S3Object[*].NAME");
let source = prepare_s3_select_statement(&mut statement).expect("terminal scalar source should be supported");
assert_eq!(source.scalar_column(), Some("name"));
assert_eq!(statement.to_string(), "SELECT NAME FROM S3Object AS NAME");
}
#[test]
fn single_quoted_source_key_adds_a_quoted_implicit_alias() {
let mut statement = parse_statement("SELECT \"Employee Data\".id FROM S3Object['Employee Data']");
let source = prepare_s3_select_statement(&mut statement).expect("single-quoted source key should be supported");
assert_eq!(
source.path(),
&[JsonPathSegment::Key {
name: "Employee Data".to_string(),
quoted: true,
}]
);
assert_eq!(source.scalar_column(), Some("Employee Data"));
assert_eq!(statement.to_string(), "SELECT \"Employee Data\".id FROM S3Object AS \"Employee Data\"");
}
#[test]
fn accepts_object_wildcard_continuation() {
let mut statement = parse_statement("SELECT * FROM S3Object[*].groups.*.id");
assert_eq!(
prepare_s3_select_statement(&mut statement)
.expect("object wildcard continuation should be supported")
.path(),
&[
JsonPathSegment::ArrayWildcard,
JsonPathSegment::Key {
name: "groups".to_string(),
quoted: false,
},
JsonPathSegment::ObjectWildcard,
JsonPathSegment::Key {
name: "id".to_string(),
quoted: false,
},
]
);
}
#[test]
fn rejects_non_literal_or_out_of_range_array_indexes() {
for sql in [
"SELECT * FROM S3Object[-1]",
"SELECT * FROM S3Object[1 + 1]",
"SELECT * FROM S3Object[999999999999999999999999999999999999]",
] {
let statement = parse_statement(sql);
assert!(
matches!(
validate_s3_select_statement(&statement),
Err(ref error)
if matches!(error.s3_select_policy_error(), Some(SelectError::DataSourcePathUnsupported))
),
"query should reject an unsafe array index: {sql}"
);
}
}
#[test]
fn accepts_group_by_and_order_by() {
let statement = parse_statement("SELECT department, COUNT(*) FROM S3Object GROUP BY department ORDER BY department");
+11 -1
View File
@@ -420,7 +420,17 @@ where
break 'updates;
};
let authoritative = match serde_json::from_slice::<DataUsageInfo>(&authoritative_data) {
Ok(info) if data_usage_info_has_persisted_baseline_identity(&info) => info,
// The bootstrap placeholder is a valid baseline identity: on a
// site that has never converged (every cycle superseded by a
// sustained write stream, #6852) it is the only authoritative
// object that will ever exist, and refusing it here means the
// observed snapshot — the only usage data such a site can
// produce — is never published at all.
Ok(info)
if data_usage_info_has_persisted_baseline_identity(&info) || data_usage_info_is_bootstrap_pending(&info) =>
{
info
}
Ok(_) => {
error!(
target: "rustfs::scanner",
@@ -30,7 +30,7 @@ catalog extension.
| `/_iceberg/v1` | Supported compatibility alias | MinIO AIStor-style alias. The smoke profile defaults to REST signing name `s3tables`. |
| S3 object data plane | Supported | Data, metadata, manifest, and delete files remain ordinary S3 objects, with table-aware policy checks for table warehouse paths. |
| Table bucket enablement | Supported | A regular RustFS bucket can be enabled for table catalog use and then addressed as the REST catalog warehouse. |
| Catalog-vended table credentials | Automated when enabled | Disabled by default. When enabled, the credentials endpoint returns short-lived table-scoped S3 credentials. |
| Catalog-vended table credentials | Automated when enabled | Disabled by default. When enabled, LoadTable vends credentials only when `X-Iceberg-Access-Delegation` contains the exact `vended-credentials` token; the dedicated credentials endpoint uses the same issuer path. |
| AWS S3 Tables endpoint shape | Profile generator | Generates the AWS catalog URI and S3 Tables warehouse ARN shape for migration docs. Full AWS S3 Tables API parity is not claimed. |
| MinIO AIStor Tables profile | Profile generator plus RustFS alias smoke | RustFS exposes the alias shape, but does not claim all AIStor private extensions. |
| Cloudflare R2 Data Catalog profile | Profile generator | Generates the catalog URI and warehouse-name shape for migration docs. Live RustFS interoperability is not claimed. |
@@ -64,12 +64,12 @@ catalog extension.
| Catalog config | Supported | `GET /v1/config` advertises RustFS catalog defaults and only the supported OpenAPI REST paths in `endpoints`. RustFS administration, maintenance, migration, diagnostics, refs, and metadata-location extensions remain available but are not presented as standard Iceberg REST endpoints. |
| Table bucket discovery | Supported | `PUT` and `GET /v1/buckets/{warehouse}` enable and inspect table bucket state. |
| Namespaces | Supported | Create, list, load, existence check, and drop namespace routes are registered on both catalog prefixes. List responses support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Namespace identifiers are limited to 512 ASCII characters so persisted paths and stateless continuation tokens remain bounded. |
| Tables | Supported | Create, register, list, load, existence check, commit, metadata-location get/update, and drop table routes are registered on both catalog prefixes. Table and view listings support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Commit identifiers must match the URL resource; unknown requirements, updates, and snapshot operations fail as bad requests; staged create, register overwrite, purge-on-drop, and v3-only encryption-key updates return an explicit unsupported-operation response. Standard statistics, partition statistics, and schema/spec cleanup updates are accepted. |
| Tables | Supported | Create, register, list, load, existence check, rename, commit, metadata-location get/update, and drop table routes are registered on both catalog prefixes. Object-backed rename uses a bucket-scoped persistent fence, recoverable intent, and conditional publication of the destination, source tombstone, and warehouse index; the source identifier is reusable only through an ETag-conditional tombstone replacement. Table and view listings support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Commit identifiers must match the URL resource; unknown requirements, updates, and snapshot operations fail as bad requests; staged create, register overwrite, purge-on-drop, and v3-only encryption-key updates return an explicit unsupported-operation response. Standard statistics, partition statistics, and schema/spec cleanup updates are accepted. |
| Commit CAS | Supported | Single-table commits validate base metadata, expected version token, referenced object existence, warehouse scope, and Iceberg commit requirements before advancing the current metadata pointer. Externally supplied metadata transitions preserve monotonic column, partition, and sequence assignment watermarks and immutable definitions for retained schemas, partition specs, sort orders, and snapshots. Standard commits preserve the normal commit-token file name and use an immutable-table-scoped fallback when rename followed by source-name reuse would otherwise collide at the same generation and commit ID. The catalog does not advertise `idempotency-key-lifetime`; clients must treat standard mutation-wide `Idempotency-Key` semantics as unsupported. |
| Commit recovery | Supported | Commit log, idempotency lookup, diagnostics, and recovery routes expose staged/finalization gaps and repair safe idempotency gaps without moving the table pointer. |
| Snapshot refs | Supported | Refs can be listed, created or replaced, and deleted through catalog commits. `main` is protected and refs with explicit retention require forced delete. |
| Iceberg views | Supported | Basic create, list, load, replace, existence check, and drop routes persist view metadata with view-scoped authorization. Replace identifiers must match the URL resource, `schema-id: -1` resolves to the last added schema, one commit timestamp is used consistently, and only Iceberg view format version 1 is accepted. |
| Table credentials endpoint | Supported | Returns an empty `storage-credentials` list by default. Returns table-scoped temporary credentials only when credential vending is enabled. Credential responses set `Cache-Control: no-store, private`, `Pragma: no-cache`, and `Expires: 0`. |
| LoadTable and table credentials endpoint | Supported | LoadTable keeps the client-provided mode unless the request negotiates `vended-credentials`. Successful vending returns one temporary session for both the table warehouse prefix and the exact current metadata location. Missing credential permission falls back to metadata-only LoadTable with an explicit reason; issuer failures remain errors. Negotiated and dedicated credential responses set `Cache-Control: no-store, private`, `Pragma: no-cache`, and `Expires: 0`. |
| Catalog diagnostics and export | Supported | Exposes recovery state, consistency state, backing manifest, recoverable commit-log WAL state, strong backing migration target, single-active-writer policy, and scale validation matrix. |
| Catalog import and rollback | Supported | Import/register and online rollback use catalog validation and commit paths rather than direct pointer mutation. Online rollback accepts only a forward-safe metadata target that preserves assignment watermarks and retained definitions. Restoring an older target that lowers those watermarks is an offline disaster-recovery operation and requires every writer to be stopped. |
| External catalog bridge | Supported operator path | Operator-supplied metadata pointer sync/import is supported for external catalog identity boundaries. Online vendor SDK polling and policy mirroring are not claimed. |
+244 -120
View File
@@ -30,73 +30,120 @@ lifetime**. Left independent, they diverge and punch through one another:
different monotonic sources, cannot be compared — a late commit fenced on one
plane can still settle quota on the other.
The fix is a single authority with one monotonic source, one persistence
The fix is a single authority with one selected comparison rule, one persistence
semantics, and one transport binding, that every consumer references rather than
re-derives.
## The authority (single source)
## Target authority and the current bounded token
**The per-object fencing epoch defined by #1312 is the sole generation
authority.** No other monotonic counter, timestamp, or random token may stand in
for generation.
The target contract still requires **one per-object commit identity** consumed
by commit fencing, read leases, cleanup, prepared reads, and quota settlement.
No consumer may mint a second value and call it the same generation.
- The distributed lock grant returns a monotonic `epoch` for the object key.
Acquiring the object write-lock is the only way to mint a new generation.
- The epoch travels down the authoritative commit path (with
`RenameDataRequest` / the local `DiskAPI` call) and is compared at each disk's
atomic `xl.meta` commit point, rejecting stale epochs. It adds no extra
network round trip (#1312 implementation clause 2).
- Every consumer in the table below **binds** this epoch. None defines its own.
The concrete ordering semantics are not settled, however. The original #1326
proposal requires a total-ordered, monotonic lock-grant epoch. Current main does
not implement that proposal. PR #6077 instead implements an opaque transaction
identity:
### Monotonicity persistence semantics
- `assign_object_transaction_epoch` mints a random non-nil UUID for PUT and
CompleteMultipartUpload when the object-transaction gate is active.
- The UUID is written through `FileInfo::set_object_transaction_epoch` into the
dual internal metadata map.
- The coordinator reads the current UUID (or `Absent`) and revalidates exact
equality immediately before `rename_data`.
- Old-data cleanup receipts carry the committed UUID and reconciliation deletes
only when the receipt UUID still equals the current object UUID.
The epoch must be **monotonic across lock-plane restart and failover**
(#1312 B4). Today the distributed lock entry is in-memory only
(`crates/lock/src/distributed_lock.rs` has no persistence path), so a lock-service
restart resets the counter to zero: a new writer draws epoch 1 while disks have
already observed epoch 100, producing either a permanent write rejection or a
fence *inversion*. To prevent this, the epoch must be one of:
This is a useful **equality-CAS fence and cleanup identity**. It is not a
monotonic epoch, is not minted by the distributed lock grant, and is not
compared atomically at each disk's `xl.meta` commit point. Until the decision
below is made, documents and issue checklists must call it the *object
transaction UUID* rather than use it as proof that the target generation
authority exists.
1. **Quorum-persisted** before it is handed to a writer, or
2. **Derived from a durable monotonic source** — a `(term, counter)` pair where
`term` advances on every lock-service leadership change and is itself durable,
so the composite never regresses even when `counter` resets.
### Ordering decision required
The comparison at the disk commit point is on the full composite; a lower
`(term, counter)` is always rejected.
Before #1313, #1314, or a unified quota binding can consume the authority, one
of these contracts must be selected and tested:
1. **Total-ordered fencing epoch.** A lock grant returns a durable per-object
`(term, counter)` (or another specified total-order type). Every disk rejects
a lower epoch at the atomic metadata commit point. The value never regresses
across lock-plane restart, failover, or minority recovery.
2. **Opaque commit-generation identity.** Consumers compare only exact identity;
no `<` / `>` semantics are permitted. The authoritative commit must perform
an atomic expected-generation CAS, and all lease, cleanup, prepared-read, and
quota contracts must be rewritten in terms of “references this exact
generation,” not “lower/newer generation.”
The current UUID implementation proves neither a durable total order nor a
per-disk atomic expected-generation CAS, so it does not by itself decide between
these options.
### Persistence semantics if total order is selected
A total-ordered epoch must be **monotonic across lock-plane restart and
failover**. The distributed lock entry remains in-memory; deriving a counter
from that entry alone would reset it after restart. The chosen source therefore
must be either quorum-persisted before grant or derived from a durable term whose
full `(term, counter)` comparison cannot regress. This requirement does not
apply to an opaque UUID as an ordering rule; the opaque alternative instead
requires atomic expected-identity comparison and durable crash recovery.
## Consumer binding contracts
### Current implementation snapshot (2026-08-31, main@9ee7b1221)
This table separates code that exists on current main from the target contract.
Closing an implementation issue does not imply that its token is already the
unified authority.
| Surface | Current main | Gap against this contract |
|---|---|---|
| PUT / CompleteMultipartUpload (#1312, PR #6077) | Owned commit tasks retain the relevant guards; an opt-in gate persists a random object transaction UUID and performs a quorum metadata equality recheck before rename | no lock-grant monotonic source; no per-disk atomic epoch/CAS comparison; the live proof is the reused remote-version-state fleet proof, not a dedicated generation capability |
| Old-data cleanup (#1323, PR #6077) | JSON receipt carries transaction UUID, old dir, and committed dir; reconciliation is gated and requires UUID equality | no generation-bound read lease is consulted, so this is crash cleanup fencing rather than the full #1313/#1323 lease lifetime contract |
| Read lease (#1313) | short-term streaming/multipart path holds the namespace read lock through EOF/drop; deterministic part-boundary coverage is tracked by PR #6887 | no cross-node generation-bound lease registry, TTL reclamation, or crash recovery |
| Prepared pool read (#1314) | PR #6889 tracks a pool-local prepared identity and fails closed/refetches when pool state changes | not merged on this snapshot; pool-local identity is not a cross-pool generation authority; black-box mixed-version/rebalance coverage remains open |
| Quota reservation (#1318) | durable per-bucket ledger plus independent snapshot-lease mutation-fence tokens; issue closed after PR #6058 | reservation and settle are not bound to the object transaction UUID; the independent fence must be reconciled with the selected authority or explicitly proven to be a separate, non-generation arbitration domain |
| Internode integrity (#1327, #1541, #1542) | v2/v3 HMAC binds audience, exact method, timestamp, nonce, canonical body digest, and receiver boot epoch; body-bound RPC policy has exact-set coverage | signature/body/replay strict switches remain default-off rollout gates; generation enforcement cannot treat an unrelated fleet-version proof as proof that these strict contracts converged |
| Consumer | How it binds generation | Key invariant |
|---|---|---|
| #1312 commit fence | epoch compared at three disk-write points — `rename`, rollback `delete`, and `commit_rename_data_dir` cleanup | stale epoch rejected on **all** disks; an already-ACK'd write is never rolled back |
| #1313 read lease | lease binds the generation observed at read time; GC runs only after every lease referencing that generation is released | lease is visible across nodes; a crashed reader's lease is reclaimed by TTL |
| #1323 old-dir GC | cleanup job carries the committed generation; before deleting `old_dir` it confirms no lease referencing a lower generation still points at it | `old_dir != committed_dir`; a still-referenced directory is never deleted |
| #1312 commit fence | selected generation is checked at `rename`, rollback restore/delete, and cleanup mutation points using the chosen ordered or exact-CAS rule | a stale writer is rejected on **all** disks; an already-ACK'd write is never rolled back |
| #1313 read lease | lease binds the exact generation observed at read time; GC runs only after every lease referencing that generation is released | lease is visible across nodes; a crashed reader's lease is reclaimed by TTL |
| #1323 old-dir GC | cleanup job carries the committed generation; before deleting `old_dir` it confirms that no lease for the generation owning that directory remains | `old_dir != committed_dir`; a still-referenced directory is never deleted |
| #1314 prepared pool read | the `PreparedPoolRead` bundle carries the generation resolved during pool lookup; the chosen pool's reader setup reuses it only after a match | generation mismatch forces a fallback to full metadata fanout |
| #1318 quota reservation | reservation / settle token binds the object generation | a late commit holding an old-generation token cannot settle a newer generation |
| #1318 quota reservation | reservation / settle record binds the exact object generation (and an ordered epoch too, if that option is selected) | a late commit cannot settle quota for a different committed generation |
### Fence coverage is three disk-write points, not one (#1312 B2)
Comparing the epoch at the `rename` commit point alone is insufficient. The
Checking the generation only before the `rename` fanout is insufficient. The
authoritative commit sequence is `tmp sync → data-dir rename → xl.meta commit →
directory sync` in `crates/ecstore/src/disk/local.rs`, and there are two further
detachable disk-write points in
`crates/ecstore/src/set_disk/core/io_primitives.rs`:
- **Rollback delete** — on quorum failure each disk runs
`delete_version(undo_write=true)`. A fenced old writer's rollback must also
compare epoch, otherwise it deletes the winner's already-committed version.
- **Rollback restore/delete** — on quorum failure each disk can restore backup
metadata or delete the failed version. A stale writer's rollback must compare
the expected generation, otherwise it can overwrite or delete the winner's
already-committed metadata.
- **`commit_rename_data_dir`** — a cancel-then-detach disk-write point; the
coordinator's "reap all child tasks" must explicitly include it so a cancelled
writer cannot bypass fence/lease and keep deleting directories.
If the epoch is validated only at the `xl.meta` commit point, a fenced writer
If generation is validated only after data-dir rename, a fenced writer
may already have renamed its data-dir into the object path, leaving a staged
orphan. Either move the fence ahead of the data-dir rename, or declare that
orphan an acceptable residue accounted for by GC metrics — the white-box
acceptance "no background disk write after release" must be rewritten
accordingly.
Current PR #6077 performs a quorum metadata equality recheck before rename and
reaps owned commit work. That closes important cancellation windows, but it is
not evidence that every disk mutation above performs the selected generation
comparison atomically. The writer inventory and per-point CAS/ordering proof
remain acceptance work for #1326 even though #1312 is closed.
### Post-commit convergence is orthogonal to the fence (#1321)
The same `SetDisks::rename_data` path already returns a post-commit
@@ -125,36 +172,40 @@ internode RPC bodies. Every such flow must be signature-bound.
### RPC signature binding (#1312 B3, #1313, #1318)
**Requirement.** The RPC body digest carrying a generation/epoch/token must be
folded into the RPC HMAC, binding `method + object key + generation`, and the
request must carry a nonce / one-shot identifier inside the 300s replay window.
The nonce is only meaningful if the **receiver enforces it**: each disk keeps a
bounded seen-nonce cache covering the 300s freshness window and rejects any
request whose nonce was already observed. A nonce that is merely transmitted but
not checked provides no replay protection.
**Requirement.** The canonical body carrying a generation or derived token must
be folded into the internode HMAC. The authenticated scope binds the target
audience, exact service/method, timestamp, nonce, canonical body digest, and
receiver replay epoch. The receiver must consume the nonce in a bounded replay
cache; transmitting a nonce without receiver-side consumption is not replay
protection.
This generalizes the existing `walk_dir` pattern: `walk_dir` computes a
`Sha256` of the request body and places it in the signed URL query as
`walk_dir_body_sha256`
(`crates/ecstore/src/cluster/rpc/internode_data_transport.rs:187`), so the body
digest is transitively covered by the URL signature. New generation-bearing RPCs
adopt the same `*_body_sha256` mechanism.
**Current substrate (verified on main).** The original legacy-only description
is obsolete:
**Current gap (verified).** The internode HMAC covers only
`{path_and_query}|{method}|{timestamp}`
(`signature_payload`, `crates/ecstore/src/cluster/rpc/http_auth.rs:75-83`). It
binds neither the request body nor a nonce, and the 300s freshness window has no
one-shot guard. Without the binding above:
- RPC v2 binds target audience, exact method, POST, timestamp, nonce, and body
digest.
- Body-bound policy covers mutating disk RPCs including `RenameData`; its
versioned canonical body includes every `RenameDataRequest` field, so the
`FileInfo` metadata map carrying the transaction UUID is authenticated.
- PR #5425 extended canonical-body enforcement to implemented non-disk mutating
unary RPCs and added an exact policy/handler coverage partition.
- PR #5455 added the receiver boot epoch and rotating replay scope so signatures
captured before a receiver restart are rejected after capability convergence.
- An on-path or replaying attacker can inject a high epoch (e.g. `u32::MAX`) and
**permanently fence out** a key's legitimate writes — monotonicity only
rejects *low/old* epochs, never a forged-high one.
- A captured lease/reservation token can be replayed within 300s to block
old-dir GC (storage-exhaustion DoS) or to double-reserve / prematurely settle
quota.
The rollout switches
`RUSTFS_INTERNODE_RPC_SIGNATURE_STRICT`,
`RUSTFS_INTERNODE_RPC_BODY_DIGEST_STRICT`, and
`RUSTFS_INTERNODE_RPC_REPLAY_SCOPE_STRICT` remain default-off for rolling
compatibility. The compatibility register and fallback/overflow metrics govern
their fleet convergence. Therefore a generation capability may claim strong
transport binding only when the relevant strict modes have converged; the
object-transaction gate's current remote-version-state fleet proof is not, by
itself, proof of RPC signature/body/replay strictness.
Acceptance for each consumer must include: "a replayed old signature to a
different method, and a forged-high-epoch request, are both rejected."
Acceptance for each generation consumer includes method substitution, canonical
body tamper, nonce replay, receiver restart, and stripped-strict-metadata
negative tests. Generation rollout must also record which strict-mode evidence
authorized enforcement.
### Encoding contract (#1312 B1)
@@ -167,11 +218,15 @@ The on-disk persistence of generation must not perturb the file format:
`xl.meta` unreadable by rolling-upgrade old RustFS nodes and by MinIO — a
total read failure, not a graceful downgrade.
- **Do not add generation as a `FileInfo` struct field.** The internode RPC layer serializes `FileInfo` with two different msgpack encoders depending on the call site: `encode_msgpack` uses rmp_serde's default **array** (positional) encoding for the `read_version` family, where a new positional field breaks decode across mixed-version nodes; `encode_msgpack_named` uses `.with_struct_map()` (named-map) encoding for `rename_data` (`crates/ecstore/src/cluster/rpc/remote_disk.rs`), which is more tolerant but still requires `#[serde(default)]` and MinIO-side agreement. Because a `FileInfo` field would have to be correct under *both* encoders and under the JSON compatibility twin (see "Wire-encoding migration" below), do not add one — use the metadata map, which rides through every encoder unchanged.
- **Where it may live.** Only inside a version's internal metadata **map**
(MinIO skips unknown internal keys and the map encoding is extensible) or in a
per-disk sidecar outside `xl.meta`. If it goes in the metadata map, it must
obey the dual-key contract (`x-rustfs-internal-*` / `x-minio-internal-*`, see
AGENTS.md "Cross-Cutting Domain Invariants").
- **Where it lives today.** The object transaction UUID uses the version's
internal metadata map under the dual-key contract
(`x-rustfs-internal-*` / `x-minio-internal-*`) via
`set_object_transaction_epoch`. Missing, malformed, nil, or conflicting dual
values fail closed when fencing is active.
- **Sidecars are not an equivalent alternative.** A future sidecar is admissible
only if it commits atomically with `xl.meta` and has a specified crash-recovery
protocol. No such protocol is implemented, so a sidecar cannot be selected by
an implementation issue merely because this document mentions one.
- **Regression guard.** Preserve the #4377 real-MinIO `xl.meta` interop
regression (the fixture family around `crates/filemeta/src/filemeta.rs`):
objects written by a new node must still be readable by old RustFS nodes and
@@ -179,82 +234,151 @@ The on-disk persistence of generation must not perturb the file format:
### Wire-encoding migration (JSON → msgpack) interaction
The internode RPC layer is mid-migration from JSON to msgpack binary, and generation-bearing fields must respect that migration window — this is not optional context, it changes how epoch is transported.
The internode RPC layer retains a JSON/msgpack rolling-compatibility window, and
generation-bearing fields must respect it.
- **Dual-field transport.** Each dual-encoded RPC field exists twice in `crates/protos/src/node.proto`: a JSON `string` field and a msgpack `bytes _bin` field (e.g. `file_info` #4 alongside `file_info_bin` #7 on `RenameDataRequest`). Senders emit both; receivers `decode_msgpack_or_json` prefer the `_bin` form and fall back to the JSON string only when `_bin` is empty (`crates/ecstore/src/cluster/rpc/remote_disk.rs`).
- **Capability flags, default off.** `rustfs_protos::internode_rpc_msgpack_only()` only drops the redundant JSON copy when both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true` are deliberately enabled after the `record_msgpack_json_fallback` metric reads zero fleet-wide and the convergence runbook is followed. If only the request flag is set, RustFS keeps dual-writing JSON compatibility fields. **Reuse this exact capability + metric-reads-zero model as the mixed-version gate for generation** rather than inventing a parallel handshake; the section above ("Capability negotiation") is layered on top of it, not instead of it.
- **Capability flags, default off.** `rustfs_protos::internode_rpc_msgpack_only()` only drops the redundant JSON copy when both `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY=true` and `RUSTFS_INTERNODE_RPC_MSGPACK_ONLY_FLEET_CONFIRMED=true` are deliberately enabled after the JSON-fallback metric reads zero fleet-wide and the convergence runbook is followed. Generation follows the same default-off, fleet-confirmed, metric-reads-zero rollout discipline, but a msgpack proof is not itself a generation capability proof.
- **Generation must ride both encodings during the window.** If epoch lives in the version's internal metadata map, that map is carried inside `FileInfo`, so it is present in both the msgpack `_bin` and JSON copies automatically — good. But any new *top-level* generation datum must be added to **both** the msgpack and JSON representations (and, for msgpack, be safe under both the array and named-map encoders). A field added to only one encoding is silently lost the moment a peer falls back to the other — exactly the failure the JSON-fallback metric exists to catch.
- **Signature must bind a canonical form.** Because a field is transmitted as both JSON and msgpack and a peer may consume either, the body-digest binding in "RPC signature binding" above must be computed over a single canonical representation (the msgpack `_bin` bytes) — not over whichever copy happened to be decoded. Once the `generation` capability is negotiated for a request, a fenced / generation-bearing request must **reject the JSON fallback path** so a downgrade to the unsigned/loosely-bound JSON copy cannot bypass the epoch check.
- **Signature binds a canonical form.** `RenameDataRequest` now has a versioned,
injective canonical-body encoder that covers both compatibility fields and is
authenticated independently of whichever JSON/msgpack decoder branch a peer
consumes. A generation-capable strict request must reject missing or
mismatched canonical-body metadata; it must not silently downgrade to an
unauthenticated JSON twin.
### Proto evolution
New generation/epoch proto fields use **proto3 `optional`** (explicit presence).
A non-optional field is forbidden: an old coordinator talking to a new disk
decodes an absent field as `0`, which is indistinguishable from a real
`epoch == 0` and silently breaks the "stale epoch rejected" invariant during
upgrade.
No top-level proto field is required by the current metadata-map UUID. If a
future ordered epoch or explicit expected-generation is added to proto, it uses
**proto3 `optional`** (explicit presence). A non-optional scalar is forbidden:
an old coordinator talking to a new disk decodes absence as a plausible zero.
### Mixed-version gate — one direction
When the cluster-level generation capability is **not** negotiated on every
target disk, the behavior **falls back to current semantics** (existing lock +
`is_lock_lost()` check for #1312; degraded-allow read-check for #1318 at
`rustfs/src/app/object/get.rs`; full fanout for #1314). Fail-closed is
**only** an explicit administrator strict mode. Defaulting to fail-closed is
forbidden — it makes writes unavailable for the whole rolling-upgrade window.
When generation enforcement is not explicitly requested, or fleet confirmation
is absent, behavior falls back to current semantics. Fail-closed is reserved for
an explicit administrator-confirmed strict rollout.
Current object transaction fencing follows that direction:
- `RUSTFS_OBJECT_TRANSACTION_FENCING_WRITE` and
`RUSTFS_OBJECT_TRANSACTION_FENCING_FLEET_CONFIRMED` both default false.
- With either flag absent, PUT/MPU does not persist or consume the transaction
UUID.
- With both flags enabled, failure to obtain or retain the live fleet proof
rejects the commit before rename.
This is an opt-in strict gate, not a negotiated generation capability. The
proof is currently borrowed from the remote-version-state writer rollout. It
proves current membership/process-epoch convergence for that feature, but does
not prove an epoch type, per-disk generation CAS support, or RPC strict-mode
convergence. Treating it as the final handshake is forbidden without an
explicit proof mapping for those properties.
## Capability negotiation
Generation enforcement is a **cluster-level handshake**, not a per-request
probe:
Generation enforcement requires one **live fleet proof**, not independent
boolean guesses in each consumer. The proof contract contains at least:
- A node advertises a `generation` capability once it can (a) mint quorum-durable
epochs, (b) compare epochs at all three disk-write points, and (c) verify the
body-digest-bound RPC signature.
- The authoritative writer enables hard enforcement for an object only when
**all** target disks in the set advertise the capability. Any missing
advertisement pins that commit to the mixed-version fallback above.
- The capability is surfaced through the existing runtime capability contract
surface (see [runtime-capability-contracts.md](runtime-capability-contracts.md)),
so consumers read one negotiated flag rather than each re-deriving support.
- Enforcement tracks the current membership rather than latching: it turns on
for a set only while every disk in that set advertises `generation`, and a
single old node rejoining drops the affected sets back to the mixed-version
fallback rather than failing closed. It never regresses the on-disk epoch —
falling back stops *comparing* new epochs, it does not lower any epoch already
persisted.
1. the selected authority version and comparison mode (ordered or exact-CAS),
2. the current membership/topology fingerprint and process epochs,
3. support for every required disk mutation point,
4. RPC signature/body/replay strict convergence, and
5. the on-disk encoding version (the current metadata-map UUID is version 1).
The authoritative writer enables enforcement only while every target disk in
the set is covered by a current proof. Membership change or an old node rejoin
revokes that proof. Revocation before commit fails an explicitly strict request;
when strict generation was never requested, the request remains on the legacy
path. Revocation never rewrites or lowers an already-persisted generation.
The existing fleet-proof machinery in `notification_sys` may be reused if its
authenticated statements are extended to cover the properties above. The
runtime capability contract may instead expose the proof. This document does
not choose the storage mechanism; it requires one token whose acquisition and
revalidation semantics are shared by all consumers.
## Implementation order
1. **#1312 first.** It defines the epoch, its persistence, the three fence
points, the RPC signature binding, and the encoding location. Everything
downstream depends on its epoch existing.
2. **#1313** (read lease) reuses the #1312 epoch as the lease generation and
must land before or alongside #1323.
3. **#1323** (old-dir GC) depends on #1313 leases being present and
cross-node-visible; its "no lease references old_dir" check has nothing to
query otherwise.
4. **#1318** (quota reservation) and **#1314** (prepared pool read) bind the
epoch independently; both gate on the same capability handshake.
Some original prerequisites have landed, but not in the originally proposed
form. Remaining work follows this order:
## Open design decisions (pin before implementation)
1. **Resolve the authority mode in #1326.** Select total order or opaque
exact-CAS, specify its atomic commit point, and audit PR #6077 against it.
Do not retrofit ordering semantics onto the existing random UUID.
2. **Define the generation fleet proof.** Map generation enablement to the RPC
signature/body/replay strict proofs delivered by #1327/#1541/#1542 and to
the selected per-disk comparison capability. Keep all strict defaults off
until fallback metrics converge.
3. **Implement #1313 generation-bound read leases.** The lease registry,
cross-node visibility, TTL, and crash recovery must exist before old-dir GC
can claim the full snapshot-lifetime guarantee. #1325 supplies the required
multi-node failure tests.
4. **Bind #1314 prepared reads.** A bundle binds the exact selected generation
within its source pool. Cross-pool ordering is forbidden until a common
authority is demonstrated. Validate rebalance and mixed-version fallback in
the #1325 multi-pool harness.
5. **Reconcile #1318 quota fencing.** Either bind reserve/settle/reconcile to
the selected object generation or document and prove that its independent
snapshot-lease fence is a separate arbitration domain that cannot settle a
different generation.
6. **Re-audit #1323 cleanup.** The existing UUID receipt remains valid crash
cleanup, but full closure against active readers requires the #1313 lease
check and the selected generation semantics.
This document fixes the transport, encoding, proto, and gate constraints, but it is not yet a complete implementable algorithm. The following must be decided and written down before any of the five consumers is coded (per the #1307 maintainer re-review, issuecomment-4992956256):
## Open design decisions (pin before contract closure)
- **Epoch type and total order.** The concrete token type and its total-order rule — a term+counter tuple, its persistence, and overflow behavior. Whether monotonicity is global or strictly per-object.
- **Never-regress on lock-service restart / minority recovery.** The epoch source must survive a lock-service restart or minority-quorum recovery without ever handing out an epoch lower than one already persisted on disk (an in-memory counter reset to zero is a fencing inversion). This is the same requirement as "Monotonicity persistence semantics" above, elevated to a hard, tested acceptance.
- **Complete xl.meta-writer coverage.** Every code path that writes xl.meta (commit rename, rollback delete/metadata restore, old-dir cleanup, heal, transition) must be enumerated and shown to compare or carry the epoch. A single unfenced writer voids the guarantee.
- **Rollback is an expected-generation CAS (#1312 B2).** The quorum-failure rollback at `io_primitives.rs:2646-2691` restores a metadata backup, not just a per-writer tmp delete, so a late rollback by writer A can overwrite writer B's committed xl.meta. Rollback must execute only when `stored_epoch == failed_writer_epoch`; a higher stored epoch must abort the rollback. Task panic / cancel / timeout at `io_primitives.rs:2602-2605` must be reaped into the coordinator's state machine, never bubble out via `?` and skip convergence.
The following decisions remain blockers for calling the contract implemented:
- **Authority mode.** Choose total order or opaque exact-CAS. If total order is
selected, define the type, per-object scope, persistence, overflow, and
never-regress restart/minority-recovery tests. If opaque identity is selected,
define the atomic expected-generation CAS and remove all ordered wording.
- **Complete xl.meta-writer coverage.** Enumerate commit rename, rollback
restore/delete, cleanup, heal, transition, restore, replication, and data
movement. Each path must compare/carry the selected generation or be proved
incapable of replacing the authoritative object identity.
- **Rollback is an expected-generation CAS (#1312 B2).** The quorum-failure
rollback in `rename_data` can restore backup metadata, not just remove a
writer-private temporary file. It must execute only when the stored generation
still matches the failed writer's expected generation. Panic, cancel, and
timeout outcomes must be reaped into coordinator convergence rather than skip
rollback through an early return.
- **Sidecar is excluded unless proven atomic.** An epoch sidecar outside `xl.meta` is only admissible if it commits at the same atomic/CAS point as `xl.meta` with a defined recovery; otherwise it opens a crash gap and must be rejected in favor of the version-internal metadata map. The earlier "metadata map or sidecar" phrasing does not treat the two as equally safe.
- **Read-lease and GC crash recovery.** Lease registry location (local vs cross-node), TTL reclamation, and crash recovery for both the lease holder and the GC executor.
- **Quota reserve → commit → settle idempotency.** The cross-stage reconcile / idempotency story for #1318, including owner-crash reconciliation, so a reservation is neither lost nor double-counted.
- **Generation capability proof.** Decide whether to extend the current
authenticated fleet proof or the runtime capability contract. It must prove
authority version, mutation coverage, topology/process epoch, and RPC strict
convergence in one revalidatable token.
- **Read-lease and GC crash recovery.** Select the cross-node registry, TTL
reclamation, lease-holder crash behavior, and GC-executor recovery. The
current cleanup receipt equality check does not answer these questions.
- **Quota reserve → commit → settle binding.** The durable ledger's idempotency
exists, but its independent mutation tokens must be related to the selected
object generation with a concrete late-settle rejection test.
- **PreparedPoolRead is pool-local only.** A #1314 bundle's generation validates freshness only within the pool that produced it. It cannot order commits across different pools unless a cross-pool common authority exists; absent that, the multi-pool wait cannot be short-circuited.
- **Hot-path cost is a blocking metric.** If per-PUT fencing grant, quota reserve, or cleanup journal adds a consensus write / fsync / centralized serialization point, it must be measured under 4KiB and high-concurrency hot-key / hot-bucket A/B as a blocking gate, not accepted by default.
- **Hot-path cost is a blocking metric.** Measure any additional consensus
write, fsync, fleet-proof lookup, lease operation, or centralized serialization
under 4 KiB and high-concurrency hot-key/hot-bucket A/B.
- **Test infrastructure.** #1325 still lacks the complete 4-node × 4-drive,
2-pool, directed network-fault, and large-object budget needed for restart,
mixed-version, and cross-node lease acceptance.
## Acceptance for this contract
- #1312 / #1313 / #1314 / #1318 / #1323 bodies reference this unified
generation and no longer define their own token.
- The five constraints — transport signature, encoding, proto presence,
mixed-version gate direction, and capability negotiation — are pinned here
once; each implementation sub-issue follows them rather than re-deciding.
- [x] Architecture document exists and is linked from the architecture index.
- [x] Transport signature, encoding, proto presence, mixed-version direction,
and capability-proof requirements are defined once.
- [x] Current implementations are separated from target guarantees; a closed
child issue is not treated as proof of unified generation binding.
- [ ] Authority mode and atomic comparison semantics are selected and tested.
- [ ] #1312 / #1313 / #1314 / #1318 / #1323 bodies reference this document and
use the selected authority terminology.
- [ ] #1313 and #1314 bind the selected generation and pass #1325 multi-node /
multi-pool failure tests.
- [ ] #1318 either binds reserve/settle to the selected generation or provides
an accepted proof that its separate fence cannot cross-settle generations.
- [ ] #1323 reconciliation checks both committed generation and active
generation-bound leases.
- [ ] Generation strict enablement is backed by one live proof that includes RPC
signature/body/replay strict convergence and per-disk comparison support.
+2 -2
View File
@@ -30,7 +30,7 @@
| chaos | 2 | |
| checksum_upload_test | 7 | |
| cluster_concurrency_test | 3 | 🌙 |
| cluster_multidrive_pool_test | 2 | 🌙 |
| cluster_multidrive_pool_test | 3 | 🌙 |
| common | 17 | |
| compression_test | 6 | ✅ |
| connection_cap_test | 2 | |
@@ -103,4 +103,4 @@
| tls_hot_reload_test | 1 | ✅ |
| version_id_regression_test | 10 | ✅ |
**Total listed: 619 tests across 86 modules · PR smoke: 165 tests / 36 modules · merge/main full: 495 tests / 77 modules · nightly replication: 56 tests · nightly cluster faults: 29 tests / 7 modules · nightly protocols: 16 tests** · updated 2026-08-26.
**Total listed: 620 tests across 86 modules · PR smoke: 165 tests / 36 modules · merge/main full: 495 tests / 77 modules · nightly replication: 56 tests · nightly cluster faults: 30 tests / 7 modules · nightly protocols: 16 tests** · updated 2026-08-31.
Generated
+6 -6
View File
@@ -2,11 +2,11 @@
"nodes": {
"nixpkgs": {
"locked": {
"lastModified": 1787209939,
"narHash": "sha256-WvvHR4kSQLAbtouMC/ruZ5UpLwlUcY3K4FAllMN+yGk=",
"lastModified": 1787964612,
"narHash": "sha256-0N9nghg3nwzX6b6qc77EzjR9cu/Z+UR66FlfsCqiURs=",
"owner": "NixOS",
"repo": "nixpkgs",
"rev": "391b592eb44808b3bd0cb80bb71b63a5a118b8bb",
"rev": "e8be7818e19ada32105a8af937a6a473b38167ca",
"type": "github"
},
"original": {
@@ -29,11 +29,11 @@
]
},
"locked": {
"lastModified": 1787454509,
"narHash": "sha256-r4LDUF+zmJnkftvCVkCrUhSJazsf6EVJF+V2l4/MYbI=",
"lastModified": 1787993548,
"narHash": "sha256-+IAEnmmx5YIhUWo0lp15jLLHchnXo5yKgWsi6C6Cf+0=",
"owner": "oxalica",
"repo": "rust-overlay",
"rev": "f60c1b57ff805a46b5175c76fc981fb4f81efbcc",
"rev": "996e9b0b019a4a9eb9e9a5641aefa06d801b5895",
"type": "github"
},
"original": {
+121 -2
View File
@@ -33,6 +33,7 @@ use super::account_audit::{
};
use super::admin_json_response;
use super::iam_error::iam_error_to_s3_error;
use super::site_replication::site_replication_iam_change_hook;
use super::supervise_admin_mutation;
use crate::admin::auth::validate_admin_request;
use crate::admin::router::{AdminOperation, Operation, S3Router};
@@ -48,6 +49,7 @@ use matchit::Params;
use rustfs_config::MAX_ADMIN_REQUEST_BODY_SIZE;
use rustfs_iam::mfa::service as mfa_service;
use rustfs_madmin::account::{AccountMfaSummary, ChangePasswordRequest, IdentityType, SelfAccountInfo, SetUserSecretKeyRequest};
use rustfs_madmin::{AccountStatus, AddOrUpdateUserReq, SITE_REPL_API_VERSION, SRIAMItem, SRIAMUser};
use rustfs_policy::auth::is_secret_key_valid;
use rustfs_policy::policy::action::{Action, AdminAction};
use rustfs_utils::MaskedAccessKey;
@@ -239,7 +241,7 @@ impl Operation for ChangeOwnPasswordHandler {
let iam_store =
current_ready_iam_handle().map_err(|_| s3::error(S3ErrorCode::InternalError, "iam is not initialized"))?;
iam_store
let (updated_at, status) = iam_store
.set_user_secret_key(&access_key, &new_secret_key)
.await
.map_err(iam_error_to_s3_error)?;
@@ -283,6 +285,11 @@ impl Operation for ChangeOwnPasswordHandler {
"admin account state"
);
// After the local revocation: peer delivery has no ordering
// dependency on it, and a slow peer must not delay killing the
// old sessions here.
broadcast_secret_key_rotation("change_own_password", &access_key, &new_secret_key, status, updated_at).await;
Ok(revoked)
})
.await?;
@@ -390,7 +397,19 @@ impl Operation for SetUserSecretKeyHandler {
let iam_store =
current_ready_iam_handle().map_err(|_| s3::error(S3ErrorCode::InternalError, "iam is not initialized"))?;
iam_store
// Derived credentials live outside the `iam-user` replication
// item: rotating one here would succeed locally and silently skip
// the peer broadcast, leaving the sites permanently diverged.
if let Some(existing) = iam_store.get_user(&target).await
&& (existing.credentials.is_temp() || existing.credentials.is_service_account())
{
return Err(s3::error(
S3ErrorCode::InvalidRequest,
"the target access key is a derived credential; rotate service accounts through update-service-account",
));
}
let (updated_at, status) = iam_store
.set_user_secret_key(&target, &request.secret_key)
.await
.map_err(iam_error_to_s3_error)?;
@@ -430,6 +449,11 @@ impl Operation for SetUserSecretKeyHandler {
"admin account state"
);
// After the local revocation: peer delivery has no ordering
// dependency on it, and a slow peer must not delay killing the
// old sessions here.
broadcast_secret_key_rotation("set_user_secret_key", &target, &request.secret_key, status, updated_at).await;
Ok(revoked)
})
.await?;
@@ -452,6 +476,58 @@ struct ChangePasswordResult {
sessions_revoked: u32,
}
/// The `iam-user` item a secret rotation fans out to peer sites.
///
/// The non-empty secret routes the peer through its create-user path (not the
/// status-only path), so the persisted status must ride along or a disabled
/// account would be re-enabled on the peer.
fn secret_key_rotation_item(access_key: &str, secret_key: &str, status: AccountStatus, updated_at: OffsetDateTime) -> SRIAMItem {
SRIAMItem {
r#type: "iam-user".to_string(),
iam_user: Some(SRIAMUser {
access_key: access_key.to_string(),
is_delete_req: false,
user_req: Some(AddOrUpdateUserReq {
secret_key: secret_key.to_string(),
policy: None,
status,
}),
api_version: Some(SITE_REPL_API_VERSION.to_string()),
}),
updated_at: Some(updated_at),
api_version: Some(SITE_REPL_API_VERSION.to_string()),
..Default::default()
}
}
/// Fan a rotated secret out to peer sites.
///
/// The rotation is already durable locally; a broadcast failure only logs,
/// matching the other IAM site-replication hooks. `status` and `updated_at`
/// come from the persisting write itself, so the item carries exactly the
/// state that was stored.
async fn broadcast_secret_key_rotation(
action: &'static str,
access_key: &str,
secret_key: &str,
status: AccountStatus,
updated_at: OffsetDateTime,
) {
if let Err(err) = site_replication_iam_change_hook(secret_key_rotation_item(access_key, secret_key, status, updated_at)).await
{
warn!(
component = LOG_COMPONENT_ADMIN,
subsystem = LOG_SUBSYSTEM_ACCOUNT,
event = EVENT_ADMIN_ACCOUNT_STATE,
action,
access_key = %MaskedAccessKey(access_key),
result = "site_replication_hook_failed",
error = ?err,
"admin account state"
);
}
}
/// Reject a new secret that would be useless or a no-op.
fn validate_new_secret_key(request: &ChangePasswordRequest) -> S3Result<()> {
if !is_secret_key_valid(&request.new_secret_key) {
@@ -499,6 +575,49 @@ mod tests {
validate_new_secret_key(&change_request("old-secret-key", "new-secret-key")).expect("must accept");
}
#[test]
fn rotation_item_takes_the_peer_create_path_and_preserves_status() {
let ts = OffsetDateTime::now_utc();
let item = secret_key_rotation_item("rotated-user", "new-secret-key", AccountStatus::Disabled, ts);
assert_eq!(item.r#type, "iam-user");
assert_eq!(item.updated_at, Some(ts));
assert!(item.api_version.is_some());
let user = item.iam_user.expect("iam-user payload");
assert_eq!(user.access_key, "rotated-user");
assert!(!user.is_delete_req);
let req = user.user_req.expect("user_req payload");
// A non-empty secret is what routes the peer through create-user
// instead of the status-only path.
assert_eq!(req.secret_key, "new-secret-key");
// Policy must stay unset so the peer's policy mapping is untouched.
assert!(req.policy.is_none());
// A disabled account must stay disabled on the peer.
assert_eq!(req.status, AccountStatus::Disabled);
}
#[test]
fn both_rotation_handlers_broadcast_after_revoking_sessions() {
let src = include_str!("account.rs");
for marker in [
"impl Operation for ChangeOwnPasswordHandler",
"impl Operation for SetUserSecretKeyHandler",
] {
let start = src.find(marker).expect("handler should exist");
let block = &src[start..];
let block = &block[..block.find("\n}\n").expect("handler block end")];
let revoke = block
.find("revoke_sts_sessions_for_parent")
.expect("handler must revoke sessions");
let broadcast = block
.find("broadcast_secret_key_rotation(")
.expect("handler must broadcast the rotation to peer sites");
assert!(revoke < broadcast, "{marker}: peer broadcast must not delay the local session revocation");
}
}
#[test]
fn route_constants_stay_under_the_admin_prefix() {
// The constants spell the full path so registration has a single source
+167 -34
View File
@@ -73,6 +73,8 @@ pub use view::*;
const JSON_CONTENT_TYPE: &str = "application/json";
const ENV_TABLE_CATALOG_CREDENTIAL_VENDING: &str = "RUSTFS_TABLE_CATALOG_CREDENTIAL_VENDING";
const ENV_TABLE_CATALOG_CREDENTIAL_TTL_SECONDS: &str = "RUSTFS_TABLE_CATALOG_CREDENTIAL_TTL_SECONDS";
const ICEBERG_ACCESS_DELEGATION_HEADER: &str = "x-iceberg-access-delegation";
const ICEBERG_VENDED_CREDENTIALS_DELEGATION: &[u8] = b"vended-credentials";
const DEFAULT_TABLE_CATALOG_CREDENTIAL_TTL_SECONDS: i64 = 15 * 60;
const MIN_TABLE_CATALOG_CREDENTIAL_TTL_SECONDS: i64 = 60;
const MAX_TABLE_CATALOG_CREDENTIAL_TTL_SECONDS: i64 = 60 * 60;
@@ -123,6 +125,7 @@ const CREDENTIAL_VENDING_UNSUPPORTED: &str = "unsupported";
const CREDENTIAL_VENDING_SUPPORTED: &str = "supported";
const CREDENTIAL_VENDING_UNSUPPORTED_REASON: &str = "temporary-credentials-not-implemented";
const CREDENTIAL_VENDING_DISABLED_REASON: &str = "credential-vending-disabled";
const CREDENTIAL_VENDING_NOT_AUTHORIZED_REASON: &str = "credential-vending-not-authorized";
const CREDENTIAL_SCOPE_WAREHOUSE_PREFIX: &str = "warehouse-prefix";
const CREDENTIAL_SCOPE_TABLE_PREFIX: &str = "table-prefix";
const CREDENTIAL_MODE_CLIENT_PROVIDED: &str = "client-provided-s3-credentials-required";
@@ -165,6 +168,7 @@ const TABLE_CATALOG_ENDPOINTS: &[&str] = &[
"GET /v1/{prefix}/namespaces/{namespace}/tables/{table}/credentials",
"POST /v1/{prefix}/namespaces/{namespace}/tables/{table}",
"DELETE /v1/{prefix}/namespaces/{namespace}/tables/{table}",
"POST /v1/{prefix}/tables/rename",
"GET /v1/{prefix}/namespaces/{namespace}/views",
"POST /v1/{prefix}/namespaces/{namespace}/views",
"GET /v1/{prefix}/namespaces/{namespace}/views/{view}",
@@ -172,10 +176,7 @@ const TABLE_CATALOG_ENDPOINTS: &[&str] = &[
"POST /v1/{prefix}/namespaces/{namespace}/views/{view}",
"DELETE /v1/{prefix}/namespaces/{namespace}/views/{view}",
];
const TABLE_CATALOG_DURABLE_STRONG_ENDPOINTS: &[&str] = &[
"POST /v1/{prefix}/namespaces/{namespace}/properties",
"POST /v1/{prefix}/tables/rename",
];
const TABLE_CATALOG_DURABLE_STRONG_ENDPOINTS: &[&str] = &["POST /v1/{prefix}/namespaces/{namespace}/properties"];
static GET_CONFIG_HANDLER: GetCatalogConfigHandler = GetCatalogConfigHandler {};
static ENABLE_TABLE_BUCKET_HANDLER: EnableTableBucketHandler = EnableTableBucketHandler {};
@@ -717,16 +718,28 @@ struct RestLoadViewResponse {
config: BTreeMap<String, String>,
}
#[derive(Debug, Serialize)]
#[derive(Serialize)]
struct RestStorageCredential {
prefix: String,
config: BTreeMap<String, String>,
}
impl std::fmt::Debug for RestStorageCredential {
fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
formatter
.debug_struct("RestStorageCredential")
.field("prefix", &self.prefix)
.field("config", &"[REDACTED]")
.finish()
}
}
#[derive(Debug, Clone)]
struct TableCredentialScope {
scope_prefix: String,
object_prefix: String,
warehouse_scope_prefix: String,
warehouse_object_prefix: String,
metadata_scope_prefix: String,
metadata_object: String,
}
#[derive(Debug, Clone)]
@@ -735,9 +748,10 @@ struct TableCredentialIssueRequest<'a> {
principal: Option<&'a rustfs_credentials::Credentials>,
scope_prefix: String,
object_prefix: String,
metadata_object: String,
}
#[derive(Debug, Clone)]
#[derive(Clone)]
struct IssuedTableCredentials {
access_key_id: String,
secret_access_key: String,
@@ -745,6 +759,18 @@ struct IssuedTableCredentials {
expiration: OffsetDateTime,
}
impl std::fmt::Debug for IssuedTableCredentials {
fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
formatter
.debug_struct("IssuedTableCredentials")
.field("access_key_id", &"[REDACTED]")
.field("secret_access_key", &"[REDACTED]")
.field("session_token", &"[REDACTED]")
.field("expiration", &self.expiration)
.finish()
}
}
#[async_trait::async_trait]
trait TableCredentialIssuer: Sync {
fn enabled(&self) -> bool {
@@ -822,7 +848,7 @@ impl TableCredentialIssuer for IamTableCredentialIssuer {
));
}
let policy = table_credential_session_policy(request.entry, &request.object_prefix)?;
let policy = table_credential_session_policy(request.entry, &request.object_prefix, &request.metadata_object)?;
let policy_buf = serde_json::to_vec(&policy)
.map_err(|err| s3_error!(InternalError, "failed to serialize table credential session policy: {}", err))?;
let expiration = OffsetDateTime::now_utc().saturating_add(Duration::seconds(self.ttl_seconds));
@@ -2270,6 +2296,7 @@ fn table_bucket_entry_from_metadata_marker(bucket: &str) -> crate::table_catalog
warehouse_root: format!("s3://{bucket}/"),
state: crate::table_catalog::TableCatalogEntryState::Active,
properties: BTreeMap::new(),
active_rename_id: None,
created_at: None,
updated_at: None,
}
@@ -2461,9 +2488,24 @@ fn table_credential_scope(entry: &crate::table_catalog::TableEntry) -> S3Result<
return Err(s3_error!(InvalidRequest, "table warehouse location must be inside the table bucket"));
}
let object_prefix = normalize_table_credential_object_prefix(object_prefix)?;
validate_persisted_table_metadata_location(entry, &entry.metadata_location)?;
let metadata_object =
crate::table_catalog::table_catalog_object_key_from_location(&entry.table_bucket, &entry.metadata_location)
.ok_or_else(|| persisted_metadata_error("table"))?;
let namespace = crate::table_catalog::Namespace::parse(&entry.namespace).map_err(|_| persisted_metadata_error("table"))?;
let table =
crate::table_catalog::IdentifierSegment::parse(entry.table.clone()).map_err(|_| persisted_metadata_error("table"))?;
if crate::table_catalog::is_reserved_table_object_key(&metadata_object)
&& !crate::table_catalog::is_valid_table_metadata_location(&namespace, &table, &metadata_object)
{
return Err(persisted_metadata_error("table"));
}
let metadata_scope_prefix = table_metadata_location_for_client(&entry.table_bucket, &entry.metadata_location);
Ok(TableCredentialScope {
scope_prefix: format!("s3://{bucket}/{object_prefix}"),
object_prefix,
warehouse_scope_prefix: format!("s3://{bucket}/{object_prefix}"),
warehouse_object_prefix: object_prefix,
metadata_scope_prefix,
metadata_object,
})
}
@@ -2501,9 +2543,16 @@ fn table_credential_catalog_resource(entry: &crate::table_catalog::TableEntry) -
Ok(format!("namespaces/{}/tables/{}", namespace.storage_id(), table.as_str()))
}
fn table_credential_session_policy(entry: &crate::table_catalog::TableEntry, object_prefix: &str) -> S3Result<Policy> {
fn table_credential_session_policy(
entry: &crate::table_catalog::TableEntry,
object_prefix: &str,
metadata_object: &str,
) -> S3Result<Policy> {
let bucket = &entry.table_bucket;
let object_prefix = normalize_table_credential_object_prefix(object_prefix)?;
validate_persisted_table_metadata_location(entry, metadata_object)?;
let metadata_object = crate::table_catalog::table_catalog_object_key_from_location(bucket, metadata_object)
.ok_or_else(|| persisted_metadata_error("table"))?;
let catalog_resource = table_credential_catalog_resource(entry)?;
let policy = serde_json::json!({
"Version": "2012-10-17",
@@ -2521,6 +2570,15 @@ fn table_credential_session_policy(entry: &crate::table_catalog::TableEntry, obj
format!("arn:aws:s3:::{bucket}/{object_prefix}*")
]
},
{
"Effect": "Allow",
"Action": [
"s3:GetObject"
],
"Resource": [
format!("arn:aws:s3:::{bucket}/{metadata_object}")
]
},
{
"Effect": "Allow",
"Action": [
@@ -2547,23 +2605,44 @@ fn table_credential_session_policy(entry: &crate::table_catalog::TableEntry, obj
Policy::parse_config(&data).map_err(|err| s3_error!(InvalidRequest, "invalid table credential policy: {}", err))
}
fn storage_credential_from_issued(scope: TableCredentialScope, issued: IssuedTableCredentials) -> RestStorageCredential {
fn storage_credential_config(scope_prefix: &str, issued: &IssuedTableCredentials) -> BTreeMap<String, String> {
let mut config = BTreeMap::new();
config.insert(S3_ACCESS_KEY_ID_CONFIG_KEY.to_string(), issued.access_key_id);
config.insert(S3_SECRET_ACCESS_KEY_CONFIG_KEY.to_string(), issued.secret_access_key);
config.insert(S3_SESSION_TOKEN_CONFIG_KEY.to_string(), issued.session_token);
config.insert(S3_ACCESS_KEY_ID_CONFIG_KEY.to_string(), issued.access_key_id.clone());
config.insert(S3_SECRET_ACCESS_KEY_CONFIG_KEY.to_string(), issued.secret_access_key.clone());
config.insert(S3_SESSION_TOKEN_CONFIG_KEY.to_string(), issued.session_token.clone());
config.insert(CREDENTIAL_VENDING_CONFIG_KEY.to_string(), CREDENTIAL_VENDING_SUPPORTED.to_string());
config.insert(CREDENTIAL_MODE_CONFIG_KEY.to_string(), CREDENTIAL_MODE_CATALOG_VENDED.to_string());
config.insert(CREDENTIAL_SCOPE_CONFIG_KEY.to_string(), CREDENTIAL_SCOPE_TABLE_PREFIX.to_string());
config.insert(CREDENTIAL_SCOPE_PREFIX_CONFIG_KEY.to_string(), scope.scope_prefix.clone());
config.insert(CREDENTIAL_SCOPE_PREFIX_CONFIG_KEY.to_string(), scope_prefix.to_string());
config.insert(
CREDENTIAL_EXPIRATION_CONFIG_KEY.to_string(),
issued.expiration.unix_timestamp().to_string(),
);
RestStorageCredential {
prefix: scope.scope_prefix,
config,
}
config
}
fn storage_credentials_from_issued(scope: TableCredentialScope, issued: IssuedTableCredentials) -> Vec<RestStorageCredential> {
let warehouse_config = storage_credential_config(&scope.warehouse_scope_prefix, &issued);
let metadata_config = storage_credential_config(&scope.metadata_scope_prefix, &issued);
vec![
RestStorageCredential {
prefix: scope.warehouse_scope_prefix,
config: warehouse_config,
},
RestStorageCredential {
prefix: scope.metadata_scope_prefix,
config: metadata_config,
},
]
}
fn requests_vended_credentials(headers: &HeaderMap) -> bool {
headers
.get_all(ICEBERG_ACCESS_DELEGATION_HEADER)
.iter()
.flat_map(|value| value.as_bytes().split(|byte| *byte == b','))
.map(|token| token.trim_ascii())
.any(|token| token == ICEBERG_VENDED_CREDENTIALS_DELEGATION)
}
fn table_metadata_location_for_client(table_bucket: &str, metadata_location: &str) -> String {
@@ -2647,6 +2726,13 @@ fn load_credentials_response_config(vending: &str, mode: &str, reason: Option<&s
config
}
fn client_provided_credentials_response(reason: &str) -> RestLoadCredentialsResponse {
RestLoadCredentialsResponse {
config: load_credentials_response_config(CREDENTIAL_VENDING_UNSUPPORTED, CREDENTIAL_MODE_CLIENT_PROVIDED, Some(reason)),
storage_credentials: Vec::new(),
}
}
fn add_table_credential_scope_config(config: &mut BTreeMap<String, String>, scope_prefix: &str) {
config.insert(CREDENTIAL_SCOPE_CONFIG_KEY.to_string(), CREDENTIAL_SCOPE_TABLE_PREFIX.to_string());
config.insert(CREDENTIAL_SCOPE_PREFIX_CONFIG_KEY.to_string(), scope_prefix.to_string());
@@ -2658,25 +2744,19 @@ async fn load_credentials_response_from_entry(
principal: Option<&rustfs_credentials::Credentials>,
) -> S3Result<RestLoadCredentialsResponse> {
if !issuer.enabled() {
return Ok(RestLoadCredentialsResponse {
config: load_credentials_response_config(
CREDENTIAL_VENDING_UNSUPPORTED,
CREDENTIAL_MODE_CLIENT_PROVIDED,
Some(CREDENTIAL_VENDING_DISABLED_REASON),
),
storage_credentials: Vec::new(),
});
return Ok(client_provided_credentials_response(CREDENTIAL_VENDING_DISABLED_REASON));
}
let scope = table_credential_scope(entry)?;
let request = TableCredentialIssueRequest {
entry,
principal,
scope_prefix: scope.scope_prefix.clone(),
object_prefix: scope.object_prefix.clone(),
scope_prefix: scope.warehouse_scope_prefix.clone(),
object_prefix: scope.warehouse_object_prefix.clone(),
metadata_object: scope.metadata_object.clone(),
};
let scope_prefix = scope.scope_prefix.clone();
let scope_prefix = scope.warehouse_scope_prefix.clone();
let storage_credentials = match issuer.issue_table_credentials(request).await? {
Some(issued) => vec![storage_credential_from_issued(scope, issued)],
Some(issued) => storage_credentials_from_issued(scope, issued),
None => {
let mut config = load_credentials_response_config(
CREDENTIAL_VENDING_UNSUPPORTED,
@@ -2698,6 +2778,28 @@ async fn load_credentials_response_from_entry(
})
}
fn apply_credentials_to_load_table_response(
mut response: RestLoadTableResponse,
credential_response: RestLoadCredentialsResponse,
) -> RestLoadTableResponse {
response.config.remove(CREDENTIAL_VENDING_CONFIG_KEY);
response.config.remove(CREDENTIAL_VENDING_REASON_CONFIG_KEY);
response.config.remove(CREDENTIAL_MODE_CONFIG_KEY);
response.config.extend(credential_response.config);
response.storage_credentials = credential_response.storage_credentials;
response
}
async fn enrich_load_table_response_with_credentials(
response: RestLoadTableResponse,
entry: &crate::table_catalog::TableEntry,
issuer: &dyn TableCredentialIssuer,
principal: Option<&rustfs_credentials::Credentials>,
) -> S3Result<RestLoadTableResponse> {
let credential_response = load_credentials_response_from_entry(entry, issuer, principal).await?;
Ok(apply_credentials_to_load_table_response(response, credential_response))
}
fn commit_table_response_from_result(
result: crate::table_catalog::TableCommitResult,
metadata: serde_json::Value,
@@ -5510,6 +5612,37 @@ async fn load_table_response<S>(
namespace: &crate::table_catalog::Namespace,
table: &str,
) -> S3Result<RestLoadTableResponse>
where
S: crate::table_catalog::TableCatalogStore + ?Sized,
{
let (entry, metadata) = load_table_entry_and_metadata(store, metadata_backend, bucket, namespace, table).await?;
Ok(load_table_response_from_entry(entry, metadata))
}
async fn load_table_response_with_credentials<S>(
store: &S,
metadata_backend: &impl crate::table_catalog::TableCatalogObjectBackend,
bucket: &str,
namespace: &crate::table_catalog::Namespace,
table: &str,
issuer: &dyn TableCredentialIssuer,
principal: Option<&rustfs_credentials::Credentials>,
) -> S3Result<RestLoadTableResponse>
where
S: crate::table_catalog::TableCatalogStore + ?Sized,
{
let (entry, metadata) = load_table_entry_and_metadata(store, metadata_backend, bucket, namespace, table).await?;
let response = load_table_response_from_entry(entry.clone(), metadata);
enrich_load_table_response_with_credentials(response, &entry, issuer, principal).await
}
async fn load_table_entry_and_metadata<S>(
store: &S,
metadata_backend: &impl crate::table_catalog::TableCatalogObjectBackend,
bucket: &str,
namespace: &crate::table_catalog::Namespace,
table: &str,
) -> S3Result<(crate::table_catalog::TableEntry, serde_json::Value)>
where
S: crate::table_catalog::TableCatalogStore + ?Sized,
{
@@ -5521,7 +5654,7 @@ where
return Err(iceberg_rest_error(ICEBERG_ERROR_NO_SUCH_TABLE, StatusCode::NOT_FOUND, "table not found"));
};
let metadata = read_persisted_table_metadata_for_entry(metadata_backend, &entry, &entry.metadata_location, true).await?;
Ok(load_table_response_from_entry(entry, metadata))
Ok((entry, metadata))
}
async fn list_views_response<S>(
@@ -121,14 +121,56 @@ impl Operation for RestLoadTableHandler {
let namespace = namespace_from_params(&params)?;
let table = table_name_from_params(&params)?;
let resource = TableCatalogResource::table(&warehouse, &namespace, &table);
authorize_table_catalog_resource_request(&req, &resource, AdminAction::GetTableMetadataAction).await?;
let principal = authorize_table_catalog_resource_request(&req, &resource, AdminAction::GetTableMetadataAction).await?;
ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?;
let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?;
let store = table_catalog_store_from_backend(metadata_backend.clone())?;
let snapshot_selection = rest_table_snapshot_selection_from_query(&req.uri)?;
let mut response = load_table_response(&store, &metadata_backend, &warehouse, &namespace, &table).await?;
let vended_credentials_requested = requests_vended_credentials(&req.headers);
let mut response = if !vended_credentials_requested {
load_table_response(&store, &metadata_backend, &warehouse, &namespace, &table).await?
} else {
let issuer = IamTableCredentialIssuer::from_request(&req)?;
let credential_permission = if issuer.enabled() {
authorize_table_catalog_resource_for_principal(
&req,
&principal,
&resource,
AdminAction::GetTableCredentialsAction,
)
.await
} else {
Ok(())
};
match credential_permission {
Ok(()) => {
load_table_response_with_credentials(
&store,
&metadata_backend,
&warehouse,
&namespace,
&table,
&issuer,
Some(&principal.credentials),
)
.await?
}
Err(err) if err.code() == &S3ErrorCode::AccessDenied => {
let response = load_table_response(&store, &metadata_backend, &warehouse, &namespace, &table).await?;
apply_credentials_to_load_table_response(
response,
client_provided_credentials_response(CREDENTIAL_VENDING_NOT_AUTHORIZED_REASON),
)
}
Err(err) => return Err(err),
}
};
apply_rest_table_snapshot_selection(&mut response.metadata, snapshot_selection);
build_json_response(StatusCode::OK, &response)
if vended_credentials_requested {
build_sensitive_json_response(StatusCode::OK, &response)
} else {
build_json_response(StatusCode::OK, &response)
}
}
}
+496 -17
View File
@@ -1,5 +1,6 @@
use super::*;
use crate::admin::runtime_sources::{AppContext, IamInterface, KmsInterface, ServerContextSlot};
use crate::storage::storage_api::contract::bucket::{BucketOperations as _, MakeBucketOptions};
use crate::table_catalog::{TableCatalogObjectBackend, TableCatalogStore};
use datafusion::{
arrow::{
@@ -43,6 +44,49 @@ impl KmsInterface for RequestKms {
}
}
fn table_catalog_handler_request(context: Arc<AppContext>, access_key: &str, secret_key: &str) -> S3Request<Body> {
let slot = ServerContextSlot::new();
assert!(slot.install(context));
let mut extensions = http::Extensions::new();
extensions.insert(slot);
S3Request {
input: Body::empty(),
method: Method::GET,
uri: "/iceberg/v1/warehouse/namespaces/analytics/tables/events"
.parse()
.expect("load table URI"),
headers: HeaderMap::new(),
extensions,
credentials: Some(s3s::auth::Credentials {
access_key: access_key.to_string(),
secret_key: s3s::auth::SecretKey::from(secret_key.to_string()),
}),
region: None,
service: None,
trailing_headers: None,
}
}
fn with_access_delegation(mut request: S3Request<Body>, values: &[&str]) -> S3Request<Body> {
for value in values {
request.headers.append(
ICEBERG_ACCESS_DELEGATION_HEADER,
HeaderValue::from_str(value).expect("access delegation header should be valid"),
);
}
request
}
async fn call_load_table_handler(request: S3Request<Body>) -> S3Result<S3Response<(StatusCode, Body)>> {
let mut router = matchit::Router::new();
router
.insert("/iceberg/v1/{warehouse}/namespaces/{namespace}/tables/{table}", ())
.expect("load table test route should insert");
let path = request.uri.path().to_string();
let matched = router.at(&path).expect("load table test route should match");
RestLoadTableHandler {}.call(request, matched.params).await
}
#[tokio::test]
async fn table_catalog_authentication_and_credentials_use_the_request_context() {
let (_temp_dir, _disk_paths, store) = crate::app::gating_test_env::isolated_multi_pool_ecstore().await;
@@ -130,6 +174,238 @@ async fn table_catalog_authentication_and_credentials_use_the_request_context()
assert!(Arc::ptr_eq(&resolved_store, &store));
}
#[tokio::test]
#[serial_test::serial]
async fn load_table_handler_negotiates_vended_credentials_without_breaking_metadata_only_callers() {
let (_temp_dir, _disk_paths, object_store) = crate::app::gating_test_env::isolated_multi_pool_ecstore().await;
object_store
.make_bucket("warehouse", &MakeBucketOptions::default())
.await
.expect("table bucket should be created");
rustfs_iam::store::object::ObjectStore::new(object_store.clone())
.save_iam_config(
serde_json::json!({"version": 1}),
format!("{}/format.json", *rustfs_iam::store::object::IAM_CONFIG_PREFIX),
)
.await
.expect("request IAM format should be seeded");
let iam = rustfs_iam::build_iam_sys(object_store.clone())
.await
.expect("request IAM should initialize");
let metadata_access_key = "load-table-metadata-only";
let metadata_secret_key = "load-table-metadata-only-secret";
let vended_access_key = "load-table-vended";
let vended_secret_key = "load-table-vended-secret";
for (access_key, secret_key) in [
(metadata_access_key, metadata_secret_key),
(vended_access_key, vended_secret_key),
] {
iam.create_user(
access_key,
&AddOrUpdateUserReq {
secret_key: secret_key.to_string(),
policy: None,
status: AccountStatus::Enabled,
},
)
.await
.expect("load table user should be created");
}
iam.set_policy(
"load-table-metadata-only-policy",
Policy::parse_config(br#"{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["admin:GetTableMetadata"]}]}"#)
.expect("metadata-only policy should parse"),
)
.await
.expect("metadata-only policy should be stored");
iam.policy_db_set(metadata_access_key, UserType::Reg, false, "load-table-metadata-only-policy")
.await
.expect("metadata-only policy should be attached");
iam.set_policy(
"load-table-vended-policy",
Policy::parse_config(br#"{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":["admin:GetTableMetadata","admin:GetTableCredentials"]}]}"#)
.expect("vended credential policy should parse"),
)
.await
.expect("vended credential policy should be stored");
iam.policy_db_set(vended_access_key, UserType::Reg, false, "load-table-vended-policy")
.await
.expect("vended credential policy should be attached");
let action_credentials = rustfs_credentials::Credentials {
access_key: "load-table-root-access-key".to_string(),
secret_key: "load-table-root-secret-key".to_string(),
status: "on".to_string(),
..Default::default()
};
let context = Arc::new(AppContext::new(
object_store.clone(),
Arc::new(RequestIam { handle: iam }),
Arc::new(RequestKms),
));
assert!(context.publish_action_credentials(action_credentials));
let setup_request = table_catalog_handler_request(context.clone(), vended_access_key, vended_secret_key);
let metadata_backend =
table_catalog_backend_from_extensions(&setup_request.extensions).expect("table catalog backend should resolve");
let catalog_store = table_catalog_store_from_backend(metadata_backend.clone()).expect("table catalog store should resolve");
enable_table_bucket_marker(object_store.as_ref(), "warehouse")
.await
.expect("table bucket should be enabled");
ensure_table_bucket_entry(&catalog_store, "warehouse", true)
.await
.expect("table bucket entry should be created");
let namespace = crate::table_catalog::Namespace::parse("analytics").expect("namespace should parse");
create_namespace_response(
&catalog_store,
"warehouse",
CreateNamespaceRequest {
namespace: vec!["analytics".to_string()],
properties: BTreeMap::new(),
},
true,
)
.await
.expect("namespace should be created");
let create_request = serde_json::from_value::<CreateTableRequest>(serde_json::json!({
"name": "events",
"schema": {
"type": "struct",
"schema-id": 0,
"fields": [{"id": 1, "name": "id", "required": true, "type": "long"}]
}
}))
.expect("create table request should parse");
create_table_response(
&catalog_store,
&TableCommitObjectBackend::trusted(metadata_backend),
"warehouse",
&namespace,
create_request,
true,
)
.await
.expect("table should be created");
let absent = temp_env::async_with_vars([(ENV_TABLE_CATALOG_CREDENTIAL_VENDING, Some("true"))], async {
call_load_table_handler(table_catalog_handler_request(context.clone(), metadata_access_key, metadata_secret_key)).await
})
.await
.expect("metadata-only caller should load without requesting delegation");
assert!(absent.headers.get(http::header::CACHE_CONTROL).is_none());
let absent_json: serde_json::Value =
serde_json::from_slice(&absent.output.1.bytes().expect("absent delegation body should be buffered"))
.expect("absent delegation response should parse");
assert_eq!(absent_json["storage-credentials"], serde_json::json!([]));
let disabled = temp_env::async_with_vars([(ENV_TABLE_CATALOG_CREDENTIAL_VENDING, None::<&str>)], async {
call_load_table_handler(with_access_delegation(
table_catalog_handler_request(context.clone(), metadata_access_key, metadata_secret_key),
&["vended-credentials"],
))
.await
})
.await
.expect("disabled vending should not add a credential permission requirement");
assert_eq!(
disabled.headers.get(http::header::CACHE_CONTROL),
Some(&HeaderValue::from_static("no-store, private"))
);
let disabled_json: serde_json::Value =
serde_json::from_slice(&disabled.output.1.bytes().expect("disabled vending body should be buffered"))
.expect("disabled vending response should parse");
assert_eq!(
disabled_json["config"][CREDENTIAL_VENDING_REASON_CONFIG_KEY],
serde_json::Value::String(CREDENTIAL_VENDING_DISABLED_REASON.to_string())
);
assert_eq!(disabled_json["storage-credentials"], serde_json::json!([]));
let remote_signing = temp_env::async_with_vars([(ENV_TABLE_CATALOG_CREDENTIAL_VENDING, Some("true"))], async {
call_load_table_handler(with_access_delegation(
table_catalog_handler_request(context.clone(), metadata_access_key, metadata_secret_key),
&["remote-signing"],
))
.await
})
.await
.expect("unrequested vending should preserve metadata-only access");
assert!(remote_signing.headers.get(http::header::CACHE_CONTROL).is_none());
let not_authorized = temp_env::async_with_vars([(ENV_TABLE_CATALOG_CREDENTIAL_VENDING, Some("true"))], async {
call_load_table_handler(with_access_delegation(
table_catalog_handler_request(context.clone(), metadata_access_key, metadata_secret_key),
&["vended-credentials"],
))
.await
})
.await
.expect("credential denial should fall back to the authorized metadata response");
let not_authorized_json: serde_json::Value = serde_json::from_slice(
&not_authorized
.output
.1
.bytes()
.expect("credential denial body should be buffered"),
)
.expect("credential denial response should parse");
assert_eq!(
not_authorized_json["config"][CREDENTIAL_VENDING_REASON_CONFIG_KEY],
serde_json::Value::String(CREDENTIAL_VENDING_NOT_AUTHORIZED_REASON.to_string())
);
assert_eq!(not_authorized_json["storage-credentials"], serde_json::json!([]));
let issued = temp_env::async_with_vars([(ENV_TABLE_CATALOG_CREDENTIAL_VENDING, Some("true"))], async {
call_load_table_handler(with_access_delegation(
table_catalog_handler_request(context.clone(), vended_access_key, vended_secret_key),
&["remote-signing", "unknown, vended-credentials"],
))
.await
})
.await
.expect("credential-authorized caller should receive vended credentials");
assert_eq!(issued.output.0, StatusCode::OK);
assert_eq!(
issued.headers.get(http::header::CACHE_CONTROL),
Some(&HeaderValue::from_static("no-store, private"))
);
assert_eq!(issued.headers.get(http::header::PRAGMA), Some(&HeaderValue::from_static("no-cache")));
assert_eq!(issued.headers.get(http::header::EXPIRES), Some(&HeaderValue::from_static("0")));
let issued_json: serde_json::Value =
serde_json::from_slice(&issued.output.1.bytes().expect("issued credential body should be buffered"))
.expect("issued credential response should parse");
assert_eq!(
issued_json["config"][CREDENTIAL_VENDING_CONFIG_KEY],
serde_json::Value::String(CREDENTIAL_VENDING_SUPPORTED.to_string())
);
assert_eq!(
issued_json["config"][CREDENTIAL_MODE_CONFIG_KEY],
serde_json::Value::String(CREDENTIAL_MODE_CATALOG_VENDED.to_string())
);
assert_eq!(issued_json["storage-credentials"].as_array().map(Vec::len), Some(2));
assert_eq!(issued_json["storage-credentials"][1]["prefix"], issued_json["metadata-location"]);
assert_eq!(
issued_json["storage-credentials"][0]["config"][S3_ACCESS_KEY_ID_CONFIG_KEY],
issued_json["storage-credentials"][1]["config"][S3_ACCESS_KEY_ID_CONFIG_KEY]
);
for required_key in [
S3_ACCESS_KEY_ID_CONFIG_KEY,
S3_SECRET_ACCESS_KEY_CONFIG_KEY,
S3_SESSION_TOKEN_CONFIG_KEY,
] {
for credential in issued_json["storage-credentials"]
.as_array()
.expect("storage credentials should be an array")
{
assert!(
credential["config"][required_key]
.as_str()
.is_some_and(|value| !value.is_empty()),
"LoadTable should include {required_key}"
);
}
}
}
#[test]
#[serial_test::serial]
fn catalog_config_response_lists_standard_rest_endpoints() {
@@ -162,7 +438,7 @@ fn catalog_config_response_lists_standard_rest_endpoints() {
.endpoints
.contains(&"POST /v1/{prefix}/namespaces/{namespace}/properties")
);
assert!(!response.endpoints.contains(&"POST /v1/{prefix}/tables/rename"));
assert!(response.endpoints.contains(&"POST /v1/{prefix}/tables/rename"));
assert_eq!(response.admin_discovery.runtime_capabilities, "/rustfs/admin/v4/runtime/capabilities");
assert_eq!(response.admin_discovery.cluster_snapshot, "/rustfs/admin/v4/cluster/snapshot");
assert_eq!(response.admin_discovery.extensions_catalog, "/rustfs/admin/v4/extensions/catalog");
@@ -9924,6 +10200,10 @@ impl TableCredentialIssuer for TestTableCredentialIssuer {
assert_eq!(request.entry.table_bucket, "warehouse");
assert_eq!(request.scope_prefix, "s3://warehouse/tables/table-id/");
assert_eq!(request.object_prefix, "tables/table-id/");
assert_eq!(
request.metadata_object,
".rustfs-table/warehouses/default/namespaces/analytics/tables/events/metadata/00001.metadata.json"
);
Ok(Some(IssuedTableCredentials {
access_key_id: "temporary-access-key".to_string(),
secret_access_key: "temporary-secret-key".to_string(),
@@ -9957,25 +10237,106 @@ async fn credential_issuer_returns_temporary_scoped_storage_credentials() {
assert!(!response.config.contains_key(S3_ACCESS_KEY_ID_CONFIG_KEY));
assert!(!response.config.contains_key(S3_SECRET_ACCESS_KEY_CONFIG_KEY));
assert!(!response.config.contains_key(S3_SESSION_TOKEN_CONFIG_KEY));
assert_eq!(response.storage_credentials.len(), 1);
let credential = &response.storage_credentials[0];
assert_eq!(credential.prefix, "s3://warehouse/tables/table-id/");
assert_eq!(credential.config.get("s3.access-key-id"), Some(&"temporary-access-key".to_string()));
assert_eq!(credential.config.get("s3.secret-access-key"), Some(&"temporary-secret-key".to_string()));
assert_eq!(credential.config.get("s3.session-token"), Some(&"temporary-session-token".to_string()));
assert_eq!(response.storage_credentials.len(), 2);
assert_eq!(response.storage_credentials[0].prefix, "s3://warehouse/tables/table-id/");
assert_eq!(
credential.config.get("rustfs.credential-mode"),
Some(&"catalog-vended-temporary-credentials".to_string())
response.storage_credentials[1].prefix,
"s3://warehouse/.rustfs-table/warehouses/default/namespaces/analytics/tables/events/metadata/00001.metadata.json"
);
for credential in &response.storage_credentials {
assert_eq!(credential.config.get("s3.access-key-id"), Some(&"temporary-access-key".to_string()));
assert_eq!(credential.config.get("s3.secret-access-key"), Some(&"temporary-secret-key".to_string()));
assert_eq!(credential.config.get("s3.session-token"), Some(&"temporary-session-token".to_string()));
assert_eq!(
credential.config.get("rustfs.credential-mode"),
Some(&"catalog-vended-temporary-credentials".to_string())
);
assert_eq!(credential.config.get("rustfs.credential-scope-prefix"), Some(&credential.prefix));
assert_eq!(
credential.config.get("rustfs.credential-expiration-unix-seconds"),
Some(&"1800000000".to_string())
);
assert!(!credential.config.contains_key("rustfs.credential-vending-reason"));
}
}
#[tokio::test]
async fn load_table_uses_the_shared_credential_vending_result() {
let entry = table_entry_for_credentials();
let metadata = serde_json::json!({
"format-version": 2,
"table-uuid": "table-uuid",
"location": "s3://warehouse/tables/table-id"
});
let principal = rustfs_credentials::Credentials {
access_key: "parent-access-key".to_string(),
secret_key: "parent-secret-key".to_string(),
..Default::default()
};
let load_table = enrich_load_table_response_with_credentials(
load_table_response_from_entry(entry.clone(), metadata),
&entry,
&TestTableCredentialIssuer,
Some(&principal),
)
.await
.expect("load table should include vended credentials");
let credentials = load_credentials_response_from_entry(&entry, &TestTableCredentialIssuer, Some(&principal))
.await
.expect("credentials endpoint should include vended credentials");
assert_eq!(
load_table.config.get(CREDENTIAL_VENDING_CONFIG_KEY),
Some(&CREDENTIAL_VENDING_SUPPORTED.to_string())
);
assert_eq!(
credential.config.get("rustfs.credential-scope-prefix"),
Some(&"s3://warehouse/tables/table-id/".to_string())
load_table.config.get(CREDENTIAL_MODE_CONFIG_KEY),
Some(&CREDENTIAL_MODE_CATALOG_VENDED.to_string())
);
assert!(!load_table.config.contains_key(CREDENTIAL_VENDING_REASON_CONFIG_KEY));
assert_eq!(
load_table.config.get(CREDENTIAL_SCOPE_PREFIX_CONFIG_KEY),
credentials.config.get(CREDENTIAL_SCOPE_PREFIX_CONFIG_KEY)
);
assert_eq!(
credential.config.get("rustfs.credential-expiration-unix-seconds"),
Some(&"1800000000".to_string())
serde_json::to_value(&load_table.storage_credentials).expect("load table credentials should serialize"),
serde_json::to_value(&credentials.storage_credentials).expect("endpoint credentials should serialize")
);
assert!(!credential.config.contains_key("rustfs.credential-vending-reason"));
}
struct RefusingTableCredentialIssuer;
#[async_trait::async_trait]
impl TableCredentialIssuer for RefusingTableCredentialIssuer {
async fn issue_table_credentials(
&self,
_request: TableCredentialIssueRequest<'_>,
) -> S3Result<Option<IssuedTableCredentials>> {
Err(S3Error::with_message(
S3ErrorCode::AccessDenied,
"table credential issuer refused request",
))
}
}
#[tokio::test]
async fn load_table_and_credentials_endpoint_propagate_issuer_refusal() {
let entry = table_entry_for_credentials();
let credentials_error = load_credentials_response_from_entry(&entry, &RefusingTableCredentialIssuer, None)
.await
.expect_err("credentials endpoint should propagate issuer refusal");
let load_table_error = enrich_load_table_response_with_credentials(
load_table_response_from_entry(entry.clone(), serde_json::json!({})),
&entry,
&RefusingTableCredentialIssuer,
None,
)
.await
.expect_err("load table should propagate issuer refusal");
assert_eq!(load_table_error.code(), credentials_error.code());
assert_eq!(load_table_error.status_code(), credentials_error.status_code());
assert_eq!(load_table_error.message(), credentials_error.message());
}
#[tokio::test]
@@ -10013,6 +10374,19 @@ fn credential_http_response_disables_caching() {
assert_eq!(response.headers.get(http::header::EXPIRES), Some(&HeaderValue::from_static("0")));
}
#[tokio::test]
async fn credential_debug_output_redacts_secrets_and_tokens() {
let response = load_credentials_response_from_entry(&table_entry_for_credentials(), &TestTableCredentialIssuer, None)
.await
.expect("issuer should build a scoped credential response");
let debug_output = format!("{response:?}");
assert!(!debug_output.contains("temporary-access-key"));
assert!(!debug_output.contains("temporary-secret-key"));
assert!(!debug_output.contains("temporary-session-token"));
assert!(debug_output.contains("[REDACTED]"));
}
#[test]
fn table_credentials_do_not_snapshot_parent_groups() {
let principal = rustfs_credentials::Credentials {
@@ -10030,7 +10404,8 @@ fn table_credentials_do_not_snapshot_parent_groups() {
#[tokio::test]
async fn table_credential_session_policy_is_limited_to_table_prefix() {
let policy = table_credential_session_policy(&table_entry_for_credentials(), "tables/table-id/")
let metadata_object = ".rustfs-table/warehouses/default/namespaces/analytics/tables/events/metadata/00001.metadata.json";
let policy = table_credential_session_policy(&table_entry_for_credentials(), "tables/table-id/", metadata_object)
.expect("table credential policy should build");
let groups = None;
let conditions = std::collections::HashMap::new();
@@ -10081,6 +10456,66 @@ async fn table_credential_session_policy_is_limited_to_table_prefix() {
})
.await
);
assert!(
policy
.is_allowed(&rustfs_policy::policy::Args {
account: "temporary-access-key",
groups: &groups,
action: Action::S3Action(rustfs_policy::policy::action::S3Action::GetObjectAction),
bucket: "warehouse",
conditions: &conditions,
is_owner: false,
object: metadata_object,
claims: &claims,
deny_only: false,
})
.await
);
assert!(
!policy
.is_allowed(&rustfs_policy::policy::Args {
account: "temporary-access-key",
groups: &groups,
action: Action::S3Action(rustfs_policy::policy::action::S3Action::PutObjectAction),
bucket: "warehouse",
conditions: &conditions,
is_owner: false,
object: metadata_object,
claims: &claims,
deny_only: false,
})
.await
);
assert!(
!policy
.is_allowed(&rustfs_policy::policy::Args {
account: "temporary-access-key",
groups: &groups,
action: Action::S3Action(rustfs_policy::policy::action::S3Action::DeleteObjectAction),
bucket: "warehouse",
conditions: &conditions,
is_owner: false,
object: metadata_object,
claims: &claims,
deny_only: false,
})
.await
);
assert!(
!policy
.is_allowed(&rustfs_policy::policy::Args {
account: "temporary-access-key",
groups: &groups,
action: Action::S3Action(rustfs_policy::policy::action::S3Action::GetObjectAction),
bucket: "warehouse",
conditions: &conditions,
is_owner: false,
object: ".rustfs-table/warehouses/default/namespaces/analytics/tables/events/metadata/00002.metadata.json",
claims: &claims,
deny_only: false,
})
.await
);
assert!(
!policy
.is_allowed(&rustfs_policy::policy::Args {
@@ -10115,8 +10550,12 @@ async fn table_credential_session_policy_is_limited_to_table_prefix() {
#[tokio::test]
async fn table_credential_session_policy_includes_table_resource_actions() {
let policy = table_credential_session_policy(&table_entry_for_credentials(), "tables/table-id/")
.expect("table credential policy should build");
let policy = table_credential_session_policy(
&table_entry_for_credentials(),
"tables/table-id/",
".rustfs-table/warehouses/default/namespaces/analytics/tables/events/metadata/00001.metadata.json",
)
.expect("table credential policy should build");
let groups = None;
let conditions = std::collections::HashMap::new();
let claims = std::collections::HashMap::new();
@@ -10177,6 +10616,45 @@ fn table_credential_scope_rejects_cross_bucket_or_unsafe_prefix() {
let mut entry = table_entry_for_credentials();
entry.warehouse_location = "s3://warehouse/tables/../table-id".to_string();
assert!(table_credential_scope(&entry).is_err());
let mut entry = table_entry_for_credentials();
entry.metadata_location = "s3://other/.rustfs-table/metadata/00001.metadata.json".to_string();
assert!(table_credential_scope(&entry).is_err());
let mut entry = table_entry_for_credentials();
entry.metadata_location =
".rustfs-table/warehouses/default/namespaces/analytics/tables/orders/metadata/00001.metadata.json".to_string();
assert!(table_credential_scope(&entry).is_err());
}
#[test]
fn table_credential_scope_accepts_entry_relative_metadata_location() {
let mut entry = table_entry_for_credentials();
entry.metadata_location = "s3://warehouse/tables/table-id/metadata/v1.metadata.json".to_string();
let scope = table_credential_scope(&entry).expect("entry-relative metadata should remain vendable");
assert_eq!(scope.metadata_object, "tables/table-id/metadata/v1.metadata.json");
assert_eq!(scope.metadata_scope_prefix, "s3://warehouse/tables/table-id/metadata/v1.metadata.json");
table_credential_session_policy(&entry, &scope.warehouse_object_prefix, &scope.metadata_object)
.expect("entry-relative metadata should produce a credential policy");
}
#[test]
fn vended_credential_delegation_requires_an_exact_comma_separated_token() {
let mut headers = HeaderMap::new();
assert!(!requests_vended_credentials(&headers));
headers.append(ICEBERG_ACCESS_DELEGATION_HEADER, HeaderValue::from_static("remote-signing"));
headers.append(ICEBERG_ACCESS_DELEGATION_HEADER, HeaderValue::from_static("unknown, vended-credentials"));
assert!(requests_vended_credentials(&headers));
let mut headers = HeaderMap::new();
headers.insert(
ICEBERG_ACCESS_DELEGATION_HEADER,
HeaderValue::from_static("not-vended-credentials, VENDED-CREDENTIALS"),
);
assert!(!requests_vended_credentials(&headers));
}
#[test]
@@ -10642,6 +11120,7 @@ async fn seed_object_table_for_metadata_maintenance(
warehouse_root: format!("s3://{bucket}/"),
state: crate::table_catalog::TableCatalogEntryState::Active,
properties: BTreeMap::new(),
active_rename_id: None,
created_at: None,
updated_at: None,
})
+92 -29
View File
@@ -21,7 +21,7 @@ use datafusion::physical_plan::SendableRecordBatchStream;
use futures::StreamExt;
use http::{HeaderMap, StatusCode, header::RANGE};
use rustfs_s3select_api::{
QueryError, SelectError,
QueryError, SelectError, SelectInputMetrics,
object_store::{INVALID_SCAN_RANGE_MESSAGE, validate_scan_range_bounds},
query::{Context, Query},
};
@@ -53,6 +53,7 @@ const UNSUPPORTED_SQL_STRUCTURE_MESSAGE: &str = "We encountered an unsupported S
struct SelectValidation {
output_format: SelectOutputFormat,
progress_enabled: bool,
reports_input_metrics: bool,
}
#[derive(Clone, Debug)]
@@ -66,6 +67,11 @@ enum SelectProducerOutcome {
ReceiverClosed,
}
struct SelectEventChannel {
tx: mpsc::Sender<S3Result<SelectObjectContentEvent>>,
terminal_permit: mpsc::OwnedPermit<S3Result<SelectObjectContentEvent>>,
}
trait SelectSnapshotFence {
fn ensure_snapshot_valid(&self) -> S3Result<()>;
}
@@ -101,10 +107,12 @@ pub async fn execute_select_object_content(
let snapshot = Arc::new(snapshot);
let query =
Query::new_with_snapshot(Context { input: input.clone() }, input.request.expression.clone(), Arc::clone(&snapshot));
let output = timeout_at(query_deadline, db.execute_admitted(&query, admission))
let query_handle = timeout_at(query_deadline, db.execute_admitted(&query, admission))
.await
.map_err(|_| select_query_timeout_error(query_timeout.as_secs()))?
.map_err(map_query_error_to_s3)?
.map_err(map_query_error_to_s3)?;
let input_metrics = Arc::clone(query_handle.query().input_metrics());
let output = query_handle
.result()
.into_record_batch_stream()
.map_err(map_query_error_to_s3)?;
@@ -121,9 +129,9 @@ pub async fn execute_select_object_content(
spawn_traced(async move {
send_select_events_until_deadline(
output,
tx,
terminal_permit,
SelectEventChannel { tx, terminal_permit },
validation,
input_metrics,
query_deadline,
query_timeout.as_secs(),
snapshot,
@@ -136,14 +144,19 @@ pub async fn execute_select_object_content(
async fn send_select_events_until_deadline<L: SelectSnapshotFence>(
output: SendableRecordBatchStream,
tx: mpsc::Sender<S3Result<SelectObjectContentEvent>>,
terminal_permit: mpsc::OwnedPermit<S3Result<SelectObjectContentEvent>>,
event_channel: SelectEventChannel,
validation: SelectValidation,
input_metrics: Arc<SelectInputMetrics>,
deadline: Instant,
timeout_seconds: u64,
snapshot_lease: L,
) {
let outcome = match timeout_at(deadline, send_select_events(output, &tx, validation, &snapshot_lease)).await {
let outcome = match timeout_at(
deadline,
send_select_events(output, &event_channel.tx, validation, input_metrics, &snapshot_lease),
)
.await
{
Ok(outcome) => outcome,
Err(_) => SelectProducerOutcome::Terminal(Err(map_query_error_to_s3(
SelectError::QueryTimeout {
@@ -153,7 +166,7 @@ async fn send_select_events_until_deadline<L: SelectSnapshotFence>(
))),
};
if let SelectProducerOutcome::Terminal(event) = outcome {
terminal_permit.send(event);
event_channel.terminal_permit.send(event);
}
drop(snapshot_lease);
}
@@ -162,10 +175,11 @@ async fn send_select_events(
mut output: SendableRecordBatchStream,
tx: &mpsc::Sender<S3Result<SelectObjectContentEvent>>,
validation: SelectValidation,
input_metrics: Arc<SelectInputMetrics>,
snapshot_fence: &impl SelectSnapshotFence,
) -> SelectProducerOutcome {
let mut encoder = SelectOutputEncoder::new(validation.output_format);
let mut progress = SelectProgress::default();
let mut progress = SelectProgress::new(validation.reports_input_metrics.then_some(input_metrics));
if tx
.send(Ok(SelectObjectContentEvent::Cont(ContinuationEvent::default())))
@@ -192,7 +206,7 @@ async fn send_select_events(
match encoder.encode_batch(&batch) {
Ok(payloads) => {
for payload in payloads {
progress.add_returned(payload.len());
let payload_len = payload.len();
if tx
.send(Ok(SelectObjectContentEvent::Records(RecordsEvent { payload: Some(payload) })))
.await
@@ -200,6 +214,7 @@ async fn send_select_events(
{
return SelectProducerOutcome::ReceiverClosed;
}
progress.add_returned(payload_len);
if validation.progress_enabled
&& tx
.send(Ok(SelectObjectContentEvent::Progress(ProgressEvent {
@@ -265,6 +280,7 @@ fn validate_select_request(headers: &http::HeaderMap, input: &mut SelectObjectCo
Ok(SelectValidation {
output_format,
progress_enabled,
reports_input_metrics: input.request.input_serialization.parquet.is_none(),
})
}
@@ -602,29 +618,39 @@ fn split_records_payload(bytes: Vec<u8>) -> Vec<Bytes> {
.collect()
}
#[derive(Default)]
struct SelectProgress {
input_metrics: Option<Arc<SelectInputMetrics>>,
bytes_returned: u64,
}
impl SelectProgress {
fn new(input_metrics: Option<Arc<SelectInputMetrics>>) -> Self {
Self {
input_metrics,
bytes_returned: 0,
}
}
fn add_returned(&mut self, bytes: usize) {
self.bytes_returned = self.bytes_returned.saturating_add(bytes as u64);
let bytes = u64::try_from(bytes).unwrap_or(u64::MAX);
self.bytes_returned = self.bytes_returned.saturating_add(bytes);
}
fn to_progress(&self) -> Progress {
let input = self.input_metrics.as_ref().map(|metrics| metrics.snapshot());
Progress {
bytes_processed: None,
bytes_processed: input.map(|metrics| clamp_i64(metrics.bytes_processed)),
bytes_returned: Some(clamp_i64(self.bytes_returned)),
bytes_scanned: None,
bytes_scanned: input.map(|metrics| clamp_i64(metrics.bytes_scanned)),
}
}
fn to_stats(&self) -> Stats {
let input = self.input_metrics.as_ref().map(|metrics| metrics.snapshot());
Stats {
bytes_processed: None,
bytes_processed: input.map(|metrics| clamp_i64(metrics.bytes_processed)),
bytes_returned: Some(clamp_i64(self.bytes_returned)),
bytes_scanned: None,
bytes_scanned: input.map(|metrics| clamp_i64(metrics.bytes_scanned)),
}
}
}
@@ -852,6 +878,7 @@ mod tests {
SelectValidation {
output_format: SelectOutputFormat::Csv(CSVOutput::default()),
progress_enabled: false,
reports_input_metrics: true,
}
}
@@ -871,9 +898,9 @@ mod tests {
let (lease, lease_released) = lease_drop_signal();
let producer = tokio::spawn(send_select_events_until_deadline(
output,
tx,
terminal_permit,
SelectEventChannel { tx, terminal_permit },
csv_validation(),
Arc::new(SelectInputMetrics::default()),
Instant::now() + std::time::Duration::from_secs(1),
300,
lease,
@@ -1096,9 +1123,9 @@ mod tests {
let (lease, lease_released) = lease_drop_signal();
let producer = tokio::spawn(send_select_events_until_deadline(
output,
tx,
terminal_permit,
SelectEventChannel { tx, terminal_permit },
csv_validation(),
Arc::new(SelectInputMetrics::default()),
Instant::now() + std::time::Duration::from_secs(1),
300,
lease,
@@ -1451,9 +1478,9 @@ mod tests {
let (lease, lease_released) = lease_drop_signal();
let producer = send_select_events_until_deadline(
output,
tx,
terminal_permit,
SelectEventChannel { tx, terminal_permit },
csv_validation(),
Arc::new(SelectInputMetrics::default()),
Instant::now() + std::time::Duration::from_secs(1),
300,
lease,
@@ -1489,7 +1516,8 @@ mod tests {
));
let (tx, mut rx) = mpsc::channel(2);
let snapshot_fence = LeaseDropSignal(None);
let producer = send_select_events(output, &tx, csv_validation(), &snapshot_fence);
let producer =
send_select_events(output, &tx, csv_validation(), Arc::new(SelectInputMetrics::default()), &snapshot_fence);
tokio::pin!(producer);
assert!(futures::poll!(producer.as_mut()).is_pending());
@@ -1525,7 +1553,14 @@ mod tests {
));
let (tx, mut rx) = mpsc::channel(4);
let outcome = send_select_events(output, &tx, csv_validation(), &FailingSnapshotFence).await;
let outcome = send_select_events(
output,
&tx,
csv_validation(),
Arc::new(SelectInputMetrics::default()),
&FailingSnapshotFence,
)
.await;
let SelectProducerOutcome::Terminal(Err(error)) = outcome else {
panic!("failed final snapshot fence must produce a terminal error");
@@ -1548,7 +1583,8 @@ mod tests {
.try_reserve_owned()
.expect("test channel should reserve terminal capacity");
let snapshot_fence = FailsAfterFirstSnapshotFence(std::sync::atomic::AtomicUsize::new(0));
let producer = send_select_events(output, &tx, csv_validation(), &snapshot_fence);
let producer =
send_select_events(output, &tx, csv_validation(), Arc::new(SelectInputMetrics::default()), &snapshot_fence);
tokio::pin!(producer);
assert!(futures::poll!(producer.as_mut()).is_pending());
@@ -1802,7 +1838,7 @@ mod tests {
#[test]
fn split_records_payload_uses_exact_returned_bytes() {
let payloads = split_records_payload(vec![b'x'; RECORDS_CHUNK_TARGET + 7]);
let mut progress = SelectProgress::default();
let mut progress = SelectProgress::new(Some(Arc::new(SelectInputMetrics::default())));
for payload in &payloads {
progress.add_returned(payload.len());
}
@@ -1922,13 +1958,40 @@ mod tests {
}
#[test]
fn progress_does_not_report_unknown_input_bytes_as_zero() {
let mut progress = SelectProgress::default();
fn progress_reports_zero_for_an_empty_input() {
let mut progress = SelectProgress::new(Some(Arc::new(SelectInputMetrics::default())));
progress.add_returned(12);
let stats = progress.to_stats();
assert_eq!(stats.bytes_returned, Some(12));
assert_eq!(stats.bytes_scanned, Some(0));
assert_eq!(stats.bytes_processed, Some(0));
}
#[test]
fn progress_dto_clamps_counters_to_signed_event_range() {
assert_eq!(clamp_i64(u64::MAX), i64::MAX);
}
#[test]
fn parquet_progress_keeps_input_metrics_unspecified() {
let mut input = base_input();
input.request.input_serialization = InputSerialization {
csv: None,
json: None,
parquet: Some(ParquetInput {}),
compression_type: None,
};
let validation = validate_select_request(&HeaderMap::new(), &mut input).expect("Parquet request should validate");
let progress = SelectProgress::new(
validation
.reports_input_metrics
.then(|| Arc::new(SelectInputMetrics::default())),
);
let stats = progress.to_stats();
assert_eq!(stats.bytes_scanned, None);
assert_eq!(stats.bytes_processed, None);
assert_eq!(stats.bytes_returned, Some(0));
}
#[test]
+8
View File
@@ -86,6 +86,14 @@ impl ApiError {
}
}
pub fn service_unavailable() -> Self {
ApiError {
code: S3ErrorCode::ServiceUnavailable,
message: Self::error_code_to_message(&S3ErrorCode::ServiceUnavailable),
source: None,
}
}
pub fn invalid_request(message: impl std::fmt::Display) -> Self {
ApiError {
code: S3ErrorCode::InvalidRequest,
+5 -1
View File
@@ -1598,7 +1598,11 @@ async fn table_data_plane_resource_for_request<T>(
error = %err,
"failed to resolve table data-plane resource"
);
s3_error!(AccessDenied, "Access Denied")
if matches!(err, crate::table_catalog::TableCatalogStoreError::Unavailable(_)) {
S3Error::from(ApiError::service_unavailable())
} else {
s3_error!(AccessDenied, "Access Denied")
}
})?;
let bucket_fence_key = (bucket.to_string(), crate::table_catalog::default_table_bucket_publication_lock_path());
let mut state = retained.state.lock();
+87 -1
View File
@@ -31,6 +31,7 @@ use crate::storage::storage_api::rpc_consumer::node_service::{
use crate::storage::storage_api::runtime_sources_consumer::{EndpointServerPools, runtime_sources};
use crate::storage::storage_api::{
sign_tonic_rpc_response_proof, verify_tonic_canonical_body_digest, verify_tonic_mutation_body_digest,
verify_tonic_mutation_body_digest_reject_unsigned,
};
use bytes::Bytes;
use futures::Stream;
@@ -123,6 +124,15 @@ fn verify_node_mutation_body<T: CanonicalMutationBody>(request: &Request<T>, ope
.map_err(|err| Status::permission_denied(format!("{operation} authentication failed: {err}")))
}
fn verify_node_signal_body<T: CanonicalMutationBody>(request: &Request<T>, operation: &'static str) -> Result<(), Status> {
let canonical_body = request
.get_ref()
.canonical_body()
.map_err(|_| Status::invalid_argument(format!("{operation} request length cannot be represented")))?;
verify_tonic_mutation_body_digest_reject_unsigned(request, &canonical_body)
.map_err(|err| Status::permission_denied(format!("{operation} authentication failed: {err}")))
}
fn start_decommission_failure_response(err: Error) -> StartDecommissionResponse {
match err {
Error::InvalidArgument(_, _, reason) => StartDecommissionResponse {
@@ -1839,7 +1849,7 @@ impl Node for NodeService {
}
async fn signal_service(&self, request: Request<SignalServiceRequest>) -> Result<Response<SignalServiceResponse>, Status> {
verify_node_mutation_body(&request, "signal service")?;
verify_node_signal_body(&request, "signal service")?;
let request = request.into_inner();
let vars = match request.vars {
Some(vars) => vars.value,
@@ -4744,6 +4754,34 @@ mod tests {
assert!(refresh_response.error_info.is_some());
}
#[tokio::test]
async fn lock_rolling_unsigned_v2_remains_compatible_for_unknown_peer() {
let service = create_test_node_service();
let unsigned_request = || {
let mut request = Request::new(GenerallyLockRequest {
args: "invalid json".to_string(),
});
request
.metadata_mut()
.insert("x-rustfs-rpc-auth-version", "2".parse().expect("valid metadata value"));
request
.metadata_mut()
.insert("x-rustfs-content-sha256", "UNSIGNED-PAYLOAD".parse().expect("valid metadata value"));
request
};
let lock = service
.lock(unsigned_request())
.await
.expect("unsigned lock must pass the rolling body gate");
assert!(!lock.into_inner().success, "invalid test lock args should fail in the lock handler");
let unlock = service
.un_lock(unsigned_request())
.await
.expect("unsigned unlock must pass the rolling body gate");
assert!(!unlock.into_inner().success, "invalid test unlock args should fail in the unlock handler");
}
/// Premise guard for the no-object-layer RPC tests (backlog#1830): they
/// assert the error surface returned while the global object layer is
/// absent. Under nextest — the authoritative runner — every test owns its
@@ -5561,6 +5599,54 @@ mod tests {
assert_eq!(response.error_info.as_deref(), Some("unsupported service signal: 99"));
}
#[tokio::test]
async fn signal_service_rejects_explicitly_unsigned_v2_body() {
let service = create_test_node_service();
let request = SignalServiceRequest {
vars: Some(Mss {
value: HashMap::from([(PEER_RESTSIGNAL.to_string(), "99".to_string())]),
}),
};
let mut request = Request::new(request);
request
.metadata_mut()
.insert("x-rustfs-rpc-auth-version", "2".parse().expect("valid metadata value"));
request
.metadata_mut()
.insert("x-rustfs-content-sha256", "UNSIGNED-PAYLOAD".parse().expect("valid metadata value"));
let error = service
.signal_service(request)
.await
.expect_err("an explicitly unsigned v2 signal must fail before handler logic");
assert_eq!(error.code(), tonic::Code::PermissionDenied);
}
#[tokio::test]
async fn signal_service_accepts_historical_unsigned_v2_marker_during_rollout() {
let service = create_test_node_service();
let mut request = Request::new(SignalServiceRequest {
vars: Some(Mss {
value: HashMap::from([(PEER_RESTSIGNAL.to_string(), "99".to_string())]),
}),
});
request
.metadata_mut()
.insert("x-rustfs-rpc-auth-version", "2".parse().expect("valid metadata value"));
request
.metadata_mut()
.insert("x-rustfs-content-sha256", "UNSIGNED-PAYLOAD".parse().expect("valid metadata value"));
request
.metadata_mut()
.insert("x-rustfs-rpc-nonce", "unsigned".parse().expect("valid metadata value"));
let response = service
.signal_service(request)
.await
.expect("historical unsigned v2 marker must remain compatible during rollout");
assert!(!response.into_inner().success, "invalid signal fixture should reach handler validation");
}
#[tokio::test]
async fn every_non_disk_mutation_rejects_a_mismatched_body_digest() {
let service = create_test_node_service();
+9 -1
View File
@@ -539,7 +539,8 @@ pub(crate) mod ecstore_rpc {
sign_ns_scanner_capability_with_tier_registry_generation, sign_put_file_capability, sign_tonic_rpc_response_proof,
tonic_boot_epoch_challenge, tonic_boot_epoch_response_headers, tonic_rpc_auth_failure_reason,
verify_put_file_auth_trailer, verify_rpc_signature, verify_tonic_canonical_body_digest,
verify_tonic_mutation_body_digest, verify_tonic_rpc_signature_with_bootstrap,
verify_tonic_mutation_body_digest, verify_tonic_mutation_body_digest_reject_unsigned,
verify_tonic_rpc_signature_with_bootstrap,
};
#[cfg(test)]
pub(crate) use rustfs_ecstore::api::rpc::{
@@ -1903,6 +1904,13 @@ pub(crate) fn verify_tonic_mutation_body_digest<T>(request: &tonic::Request<T>,
ecstore_rpc::verify_tonic_mutation_body_digest(request, canonical_body)
}
pub(crate) fn verify_tonic_mutation_body_digest_reject_unsigned<T>(
request: &tonic::Request<T>,
canonical_body: &[u8],
) -> std::io::Result<()> {
ecstore_rpc::verify_tonic_mutation_body_digest_reject_unsigned(request, canonical_body)
}
#[cfg(test)]
pub(crate) fn set_tonic_canonical_body_digest<T>(request: &mut tonic::Request<T>, canonical_body: &[u8]) -> std::io::Result<()> {
ecstore_rpc::set_tonic_canonical_body_digest(request, canonical_body)
+2
View File
@@ -101,6 +101,7 @@ pub(crate) const TABLE_RESOURCE_MARKER_VERSION: u16 = 1;
)]
pub(crate) const TABLE_METADATA_POINTER_VERSION: u16 = 1;
pub(crate) const TABLE_CATALOG_ENTRY_VERSION: u16 = 1;
pub(crate) const TABLE_RENAME_INTENT_VERSION: u16 = 1;
pub(crate) const TABLE_WAREHOUSE_INDEX_STATE_VERSION: u16 = 2;
pub(crate) const TABLE_MAINTENANCE_CONFIG_VERSION: u16 = 1;
pub(crate) const TABLE_EXTERNAL_CATALOG_BRIDGE_VERSION: u16 = 1;
@@ -166,6 +167,7 @@ const COMMIT_LOG_ROOT: &str = "commits";
const COMMIT_IDEMPOTENCY_ROOT: &str = "commit-idempotency";
const WAREHOUSE_INDEX_ROOT: &str = "warehouse-index";
const WAREHOUSE_INDEX_STATE_FILE: &str = "state.json";
const TABLE_RENAME_ROOT: &str = "renames";
const WAREHOUSE_INDEX_MAX_PREFIX_DEPTH: usize = 64;
const EXTERNAL_CATALOG_ROOT: &str = "external-catalog";
const EXTERNAL_CATALOG_BRIDGE_FILE: &str = "bridge.json";
+33
View File
@@ -39,6 +39,8 @@ pub(crate) fn table_bucket_marker_json() -> Result<Vec<u8>, serde_json::Error> {
#[serde(rename_all = "SCREAMING_SNAKE_CASE")]
pub(crate) enum TableCatalogEntryState {
Active,
/// Persisted only behind a rename intent so older readers reject the unknown state and fail closed.
Renaming,
Deleting,
Deleted,
}
@@ -53,10 +55,40 @@ pub(crate) struct TableBucketEntry {
pub state: TableCatalogEntryState,
#[serde(default)]
pub properties: BTreeMap<String, String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub active_rename_id: Option<String>,
pub created_at: Option<String>,
pub updated_at: Option<String>,
}
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
#[serde(rename_all = "SCREAMING_SNAKE_CASE")]
pub(crate) enum TableRenameIntentState {
Prepared,
SourceFenced,
DestinationWritten,
SourceTombstoned,
IndexPublished,
DestinationPublished,
Completed,
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub(crate) struct TableRenameIntent {
pub version: u16,
pub rename_id: String,
pub table_bucket: String,
pub source: TableEntry,
pub destination: TableEntry,
pub source_etag: String,
pub destination_etag: Option<String>,
pub warehouse_index_etag: String,
pub state: TableRenameIntentState,
pub created_at: String,
pub updated_at: String,
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub(crate) struct NamespaceEntry {
@@ -1225,6 +1257,7 @@ pub(crate) enum TableCatalogBackingMigrationStep {
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
#[serde(rename_all = "SCREAMING_SNAKE_CASE")]
pub(crate) enum TableCatalogBackingMigrationBlocker {
TableRenameRecoveryRequired,
CommitRecoveryRequired,
CommitManualReviewRequired,
WarehouseIndexBackfillRequired,
+33 -15
View File
@@ -301,6 +301,11 @@ where
return Err(TableCatalogStoreError::NotFound(format!("table bucket {table_bucket}")));
};
validate_table_bucket_entry_object(&self.paths, &bucket_path, &table_bucket_entry)?;
if table_bucket_entry.active_rename_id.is_some() {
return Err(TableCatalogStoreError::Conflict(format!(
"table bucket {table_bucket} has a table rename requiring recovery"
)));
}
let mut namespaces = Vec::new();
let mut tables = Vec::new();
@@ -343,6 +348,9 @@ where
)));
};
validate_table_entry_object(&self.paths, table_object, &table_entry)?;
if table_entry.state != TableCatalogEntryState::Active {
continue;
}
for commit_object in self
.backend
@@ -595,9 +603,10 @@ where
&self,
table_bucket: &str,
) -> TableCatalogStoreResult<TableCatalogBackingMigrationDryRunReport> {
if self.get_table_bucket(table_bucket).await?.is_none() {
let Some(table_bucket_entry) = self.get_table_bucket(table_bucket).await? else {
return Err(TableCatalogStoreError::NotFound(format!("table bucket {table_bucket}")));
}
};
let rename_recovery_required = table_bucket_entry.active_rename_id.is_some();
if let Some((global_fence, _)) = self
.read_entry::<TableCatalogBackingMigrationGlobalFence>(
self.catalog_bucket(),
@@ -653,18 +662,19 @@ where
continue;
};
validate_table_entry_object(&self.paths, &object, &table)?;
if table.state != TableCatalogEntryState::Active {
continue;
}
table_count = table_count.saturating_add(1);
if !table_ids.insert(table.table_id.clone()) {
duplicate_table_identity = true;
}
if table.state == TableCatalogEntryState::Active {
active_table_identifiers.insert((table.namespace.clone(), table.table.clone()));
let warehouse_prefix = table_warehouse_object_prefix(&table)?;
warehouse_prefix_owners
.entry(warehouse_prefix)
.and_modify(|count| *count = count.saturating_add(1))
.or_insert(1);
}
active_table_identifiers.insert((table.namespace.clone(), table.table.clone()));
let warehouse_prefix = table_warehouse_object_prefix(&table)?;
warehouse_prefix_owners
.entry(warehouse_prefix)
.and_modify(|count| *count = count.saturating_add(1))
.or_insert(1);
let recovery = self.table_commit_recovery_report_for_entry(&table, 0).await?;
commit_log_count = commit_log_count.saturating_add(recovery.commits.len());
@@ -711,33 +721,41 @@ where
let table_view_identifier_collision_count = active_table_identifiers.intersection(&active_view_identifiers).count();
let mut blockers = Vec::new();
let mut recommended_actions = Vec::new();
if rename_recovery_required {
blockers.push(TableCatalogBackingMigrationBlocker::TableRenameRecoveryRequired);
recommended_actions.push(TableCatalogBackingMigrationAction::RunCatalogRecovery);
}
if recovery_required_count > 0 {
blockers.push(TableCatalogBackingMigrationBlocker::CommitRecoveryRequired);
}
if manual_review_count > 0 {
blockers.push(TableCatalogBackingMigrationBlocker::CommitManualReviewRequired);
}
if recovery_required_count > 0 || manual_review_count > 0 {
if (recovery_required_count > 0 || manual_review_count > 0)
&& !recommended_actions.contains(&TableCatalogBackingMigrationAction::RunCatalogRecovery)
{
recommended_actions.push(TableCatalogBackingMigrationAction::RunCatalogRecovery);
}
if !warehouse_index_ready {
blockers.push(TableCatalogBackingMigrationBlocker::WarehouseIndexBackfillRequired);
recommended_actions.push(TableCatalogBackingMigrationAction::BackfillWarehouseIndex);
}
if conflicting_warehouse_prefix {
if conflicting_warehouse_prefix && !rename_recovery_required {
blockers.push(TableCatalogBackingMigrationBlocker::DuplicateWarehousePrefix);
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewDuplicateWarehousePrefixes);
}
if duplicate_table_identity {
if duplicate_table_identity && !rename_recovery_required {
blockers.push(TableCatalogBackingMigrationBlocker::DuplicateTableIdentity);
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewDuplicateTableIdentities);
}
if table_view_identifier_collision_count > 0 {
if table_view_identifier_collision_count > 0 && !rename_recovery_required {
blockers.push(TableCatalogBackingMigrationBlocker::TableViewIdentifierCollision);
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewTableViewIdentifierCollisions);
}
let mut status = if manual_review_count > 0
let mut status = if rename_recovery_required {
TableCatalogBackingMigrationStatus::RecoveryRequired
} else if manual_review_count > 0
|| conflicting_warehouse_prefix
|| duplicate_table_identity
|| table_view_identifier_collision_count > 0
+17 -3
View File
@@ -57,6 +57,9 @@ fn validate_table_bucket_entry(entry: &TableBucketEntry) -> TableCatalogStoreRes
if entry.catalog_type != TABLE_BUCKET_CATALOG_TYPE {
return Err(TableCatalogStoreError::Invalid("unsupported table bucket catalog type".to_string()));
}
if entry.active_rename_id.as_ref().is_some_and(String::is_empty) {
return Err(TableCatalogStoreError::Invalid("active table rename id cannot be empty".to_string()));
}
Ok(())
}
@@ -984,6 +987,15 @@ impl TableCatalogObjectPaths {
)
}
pub fn table_rename_intent_path(&self, table_bucket: &str, rename_id: &str) -> String {
format!(
"{}{}/{}.json",
self.table_bucket_root_prefix(table_bucket),
TABLE_RENAME_ROOT,
table_catalog_path_hash(rename_id)
)
}
pub fn backing_migration_fence_path(&self, table_bucket: &str) -> String {
format!(
"{}{}/{}",
@@ -1241,9 +1253,11 @@ where
destination_table: &str,
) -> TableCatalogStoreResult<()> {
match self {
Self::ObjectBacked(_) => Err(TableCatalogStoreError::Unsupported(
"table rename requires durable-strong catalog backing".to_string(),
)),
Self::ObjectBacked(store) => {
store
.rename_table(table_bucket, source_namespace, source_table, destination_namespace, destination_table)
.await
}
Self::DurableStrong(store) => {
store
.rename_table(table_bucket, source_namespace, source_table, destination_namespace, destination_table)
File diff suppressed because it is too large Load Diff
+15 -1
View File
@@ -642,12 +642,26 @@ impl TestCatalogObjectBackend {
let mut state = self.state.lock().await;
let key = (bucket.to_string(), object.to_string());
let next_attempt = state.put_attempts.get(&key).copied().unwrap_or_default() + 1;
Self::pause_put_attempt_unlocked(&mut state, key, next_attempt)
}
pub(crate) async fn pause_put_attempt(&self, bucket: &str, object: &str, attempt: usize) -> TestCatalogObjectPause {
let mut state = self.state.lock().await;
let key = (bucket.to_string(), object.to_string());
Self::pause_put_attempt_unlocked(&mut state, key, attempt)
}
fn pause_put_attempt_unlocked(
state: &mut TestCatalogObjectState,
key: (String, String),
attempt: usize,
) -> TestCatalogObjectPause {
let pause = TestCatalogObjectPause::default();
state
.pause_put_attempts
.entry(key)
.or_default()
.insert(next_attempt, pause.clone());
.insert(attempt, pause.clone());
pause
}
+460 -3
View File
@@ -144,6 +144,7 @@ fn catalog_entry_structures_serialize_stable_fields() {
warehouse_root: "s3://analytics/".to_string(),
state: TableCatalogEntryState::Active,
properties: BTreeMap::from([("owner".to_string(), "platform".to_string())]),
active_rename_id: None,
created_at: Some("2026-05-23T00:00:00Z".to_string()),
updated_at: Some("2026-05-23T00:00:00Z".to_string()),
};
@@ -3441,6 +3442,7 @@ fn test_bucket_entry(bucket: &str) -> TableBucketEntry {
warehouse_root: format!("s3://{bucket}/"),
state: TableCatalogEntryState::Active,
properties: BTreeMap::new(),
active_rename_id: None,
created_at: None,
updated_at: None,
}
@@ -5107,7 +5109,8 @@ async fn object_catalog_pagination_bounds_reads_and_covers_rest_resources() {
.await
.expect("first table page should load");
assert_eq!(table_page.entries[0].table, "alpha");
assert_eq!(backend.read_call_count().await, 1);
// One read snapshots the bucket rename fence and one loads the page entry.
assert_eq!(backend.read_call_count().await, 2);
let table_page = store
.list_tables_page(bucket, &namespace_name, table_page.next_cursor.as_deref(), one)
.await
@@ -17752,7 +17755,461 @@ async fn strong_catalog_table_rename_returns_success_after_committed_snapshot_re
}
#[tokio::test]
async fn configured_object_catalog_rejects_table_rename() {
async fn object_catalog_table_rename_preserves_identity_index_and_reuses_source_tombstone() {
let backend = TestCatalogObjectBackend {
content_addressed_etags: true,
..Default::default()
};
let store = ObjectTableCatalogStore::new(backend);
let bucket = "analytics";
let source_namespace = Namespace::parse("sales").expect("source namespace should parse");
let destination_namespace = Namespace::parse("curated").expect("destination namespace should parse");
let source_table = IdentifierSegment::parse("orders").expect("source table should parse");
store.put_table_bucket(test_bucket_entry(bucket)).await.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &source_namespace))
.await
.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &destination_namespace))
.await
.unwrap();
let source = test_table_entry(
bucket,
&source_namespace,
&source_table,
default_table_metadata_file_path(&source_namespace, &source_table, "00001.metadata.json"),
);
store.create_table(source.clone()).await.unwrap();
let bucket_object = store.paths.table_bucket_entry_path(bucket);
let bucket_etag_before = store
.read_entry::<TableBucketEntry>(RUSTFS_META_BUCKET, &bucket_object)
.await
.unwrap()
.unwrap()
.1
.expect("table bucket should have an etag");
store
.rename_table(bucket, "sales", "orders", "curated", "orders_v2")
.await
.expect("object-backed table rename should complete");
let bucket_etag_after = store
.read_entry::<TableBucketEntry>(RUSTFS_META_BUCKET, &bucket_object)
.await
.unwrap()
.unwrap()
.1
.expect("table bucket should have an etag");
assert_ne!(bucket_etag_after, bucket_etag_before);
assert!(store.load_table(bucket, "sales", "orders").await.unwrap().is_none());
let destination = store
.load_table(bucket, "curated", "orders_v2")
.await
.unwrap()
.expect("destination table should exist");
let mut expected_destination = source.clone();
expected_destination.namespace = "curated".to_string();
expected_destination.table = "orders_v2".to_string();
assert_ne!(destination.updated_at, source.updated_at);
expected_destination.updated_at.clone_from(&destination.updated_at);
assert_eq!(destination, expected_destination);
let source_object = store.paths.table_entry_path(bucket, &source_namespace, &source_table);
let source_tombstone = store
.read_entry::<TableEntry>(RUSTFS_META_BUCKET, &source_object)
.await
.unwrap()
.expect("source tombstone should remain")
.0;
assert_eq!(source_tombstone.state, TableCatalogEntryState::Deleted);
assert_eq!(source_tombstone.table_id, destination.table_id);
assert!(
store
.get_table_bucket(bucket)
.await
.unwrap()
.expect("table bucket should exist")
.active_rename_id
.is_none()
);
let destination_index = table_warehouse_index_entry(&destination).unwrap();
let index_object = store
.paths
.warehouse_index_entry_path(bucket, &destination_index.warehouse_object_prefix);
let persisted_index = store
.read_entry::<TableWarehouseIndexEntry>(RUSTFS_META_BUCKET, &index_object)
.await
.unwrap()
.expect("warehouse index should exist")
.0;
assert_eq!(persisted_index, destination_index);
let resource = store
.resolve_table_data_plane_resource(bucket, "tables/table-id/data/part.parquet")
.await
.unwrap()
.expect("renamed table should resolve its stable warehouse prefix");
assert_eq!(resource.namespace, "curated");
assert_eq!(resource.table, "orders_v2");
store
.backfill_table_warehouse_index(bucket)
.await
.expect("retained source tombstone should not make the warehouse index ambiguous");
let migration = store.plan_durable_strong_backing_migration(bucket).await.unwrap();
assert_eq!(migration.table_count, 1);
assert!(
!migration
.blockers
.contains(&TableCatalogBackingMigrationBlocker::DuplicateTableIdentity)
);
store
.rename_table(bucket, "curated", "orders_v2", "sales", "orders")
.await
.expect("rename should conditionally replace the retained source tombstone");
store
.rename_table(bucket, "sales", "orders", "curated", "orders_v2")
.await
.expect("rename should conditionally replace a destination tombstone");
let mut replacement = test_table_entry(
bucket,
&source_namespace,
&source_table,
default_table_metadata_file_path(&source_namespace, &source_table, "00002.metadata.json"),
);
replacement.table_id = "replacement-table-id".to_string();
replacement.table_uuid = "replacement-table-uuid".to_string();
replacement.warehouse_location = "s3://analytics/tables/replacement-table-id".to_string();
store
.create_table(replacement.clone())
.await
.expect("create should conditionally replace the source tombstone");
assert_eq!(
store
.load_table(bucket, "sales", "orders")
.await
.unwrap()
.expect("source identifier should be reusable"),
replacement
);
}
#[tokio::test]
async fn object_catalog_table_rename_rejects_missing_and_conflicting_destinations() {
let backend = TestCatalogObjectBackend::default();
let store = ObjectTableCatalogStore::new(backend);
let bucket = "analytics";
let source_namespace = Namespace::parse("sales").unwrap();
let destination_namespace = Namespace::parse("curated").unwrap();
let source_table = IdentifierSegment::parse("orders").unwrap();
let destination_table = IdentifierSegment::parse("orders_v2").unwrap();
store.put_table_bucket(test_bucket_entry(bucket)).await.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &source_namespace))
.await
.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &destination_namespace))
.await
.unwrap();
store
.create_table(test_table_entry(
bucket,
&source_namespace,
&source_table,
default_table_metadata_file_path(&source_namespace, &source_table, "00001.metadata.json"),
))
.await
.unwrap();
assert_matches!(
store.rename_table(bucket, "sales", "orders", "missing", "orders_v2").await,
Err(TableCatalogStoreError::NamespaceNotFound(_))
);
assert_matches!(
store.rename_table(bucket, "sales", "missing", "curated", "orders_v2").await,
Err(TableCatalogStoreError::TableNotFound(_))
);
let mut existing = test_table_entry(
bucket,
&destination_namespace,
&destination_table,
default_table_metadata_file_path(&destination_namespace, &destination_table, "00001.metadata.json"),
);
existing.table_id = "destination-table-id".to_string();
existing.table_uuid = "destination-table-uuid".to_string();
existing.warehouse_location = "s3://analytics/tables/destination-table-id".to_string();
store.create_table(existing).await.unwrap();
assert_matches!(
store.rename_table(bucket, "sales", "orders", "curated", "orders_v2").await,
Err(TableCatalogStoreError::AlreadyExists(_))
);
let destination_view = IdentifierSegment::parse("orders_view").unwrap();
store
.create_view(test_view_entry(
bucket,
&destination_namespace,
&destination_view,
default_view_metadata_file_path(&destination_namespace, &destination_view, "00001.metadata.json"),
))
.await
.unwrap();
assert_matches!(
store.rename_table(bucket, "sales", "orders", "curated", "orders_view").await,
Err(TableCatalogStoreError::AlreadyExists(_))
);
assert!(store.load_table(bucket, "sales", "orders").await.unwrap().is_some());
}
#[tokio::test]
async fn object_catalog_table_rename_fails_closed_around_durable_fence_creation() {
let backend = TestCatalogObjectBackend::default();
let store = ObjectTableCatalogStore::new(backend.clone());
let bucket = "analytics";
let source_namespace = Namespace::parse("sales").unwrap();
let destination_namespace = Namespace::parse("curated").unwrap();
let source_table = IdentifierSegment::parse("orders").unwrap();
store.put_table_bucket(test_bucket_entry(bucket)).await.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &source_namespace))
.await
.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &destination_namespace))
.await
.unwrap();
store
.create_table(test_table_entry(
bucket,
&source_namespace,
&source_table,
default_table_metadata_file_path(&source_namespace, &source_table, "00001.metadata.json"),
))
.await
.unwrap();
let bucket_object = store.paths.table_bucket_entry_path(bucket);
backend.fail_next_put(RUSTFS_META_BUCKET, &bucket_object).await;
assert_matches!(
store.rename_table(bucket, "sales", "orders", "curated", "orders_v2").await,
Err(TableCatalogStoreError::Internal(_))
);
assert!(
store
.get_table_bucket(bucket)
.await
.unwrap()
.expect("table bucket should remain")
.active_rename_id
.is_none()
);
assert!(store.load_table(bucket, "sales", "orders").await.unwrap().is_some());
assert!(store.load_table(bucket, "curated", "orders_v2").await.unwrap().is_none());
let mut fenced_bucket = store.get_table_bucket(bucket).await.unwrap().unwrap();
fenced_bucket.active_rename_id = Some("missing-intent".to_string());
backend
.seed_object(
RUSTFS_META_BUCKET,
&bucket_object,
serde_json::to_vec(&fenced_bucket).expect("fenced bucket should serialize"),
)
.await;
assert_matches!(
store.rename_table(bucket, "sales", "orders", "curated", "orders_v2").await,
Err(TableCatalogStoreError::Unavailable(_))
);
assert_eq!(
store
.get_table_bucket(bucket)
.await
.unwrap()
.expect("table bucket should remain fail-closed")
.active_rename_id
.as_deref(),
Some("missing-intent")
);
assert_matches!(
store
.resolve_table_data_plane_resource(bucket, "tables/table-id/data/part.parquet")
.await,
Err(TableCatalogStoreError::Unavailable(_))
);
}
#[tokio::test]
async fn object_catalog_table_rename_recovers_after_destination_publish_and_fences_concurrent_mutations() {
let backend = TestCatalogObjectBackend::default();
let store = ObjectTableCatalogStore::new(backend.clone());
let bucket = "analytics";
let source_namespace = Namespace::parse("sales").expect("source namespace should parse");
let destination_namespace = Namespace::parse("curated").expect("destination namespace should parse");
let source_table = IdentifierSegment::parse("orders").expect("source table should parse");
store.put_table_bucket(test_bucket_entry(bucket)).await.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &source_namespace))
.await
.unwrap();
store
.create_namespace(test_namespace_entry(bucket, &destination_namespace))
.await
.unwrap();
let source = test_table_entry(
bucket,
&source_namespace,
&source_table,
default_table_metadata_file_path(&source_namespace, &source_table, "00001.metadata.json"),
);
store.create_table(source.clone()).await.unwrap();
let source_object = store.paths.table_entry_path(bucket, &source_namespace, &source_table);
let source_tombstone_attempt = backend.put_attempt_count(RUSTFS_META_BUCKET, &source_object).await + 2;
let source_tombstone_pause = backend
.pause_put_attempt(RUSTFS_META_BUCKET, &source_object, source_tombstone_attempt)
.await;
let rename_store = store.clone();
let rename = tokio::spawn(async move {
rename_store
.rename_table(bucket, "sales", "orders", "curated", "orders_v2")
.await
});
source_tombstone_pause.wait_started().await;
let active_rename_id = store
.get_table_bucket(bucket)
.await
.unwrap()
.expect("table bucket should exist")
.active_rename_id
.expect("rename fence should be durable before destination publication");
assert_matches!(
store.load_table(bucket, "sales", "orders").await,
Err(TableCatalogStoreError::Unavailable(_))
);
assert_matches!(store.list_tables(bucket, "sales").await, Err(TableCatalogStoreError::Unavailable(_)));
assert_matches!(
store
.resolve_table_data_plane_resource(bucket, "tables/table-id/data/part.parquet")
.await,
Err(TableCatalogStoreError::Unavailable(_))
);
let migration = store.plan_durable_strong_backing_migration(bucket).await.unwrap();
assert_eq!(migration.status, TableCatalogBackingMigrationStatus::RecoveryRequired);
assert!(
migration
.blockers
.contains(&TableCatalogBackingMigrationBlocker::TableRenameRecoveryRequired)
);
let publication_lock = default_table_bucket_publication_lock_path();
let publication_attempts = backend.write_lock_acquisition_count(bucket, &publication_lock).await;
let commit_store = store.clone();
let commit = tokio::spawn(async move {
commit_store
.commit_table(TableCommitRequest {
table_bucket: bucket.to_string(),
namespace: "sales".to_string(),
table: "orders".to_string(),
commit_id: "concurrent-commit".to_string(),
idempotency_key: None,
operation: "append".to_string(),
expected_version_token: source.version_token,
expected_metadata_location: source.metadata_location,
new_metadata_location: "unused.metadata.json".to_string(),
requirements: Vec::new(),
writer: Some("rename-test".to_string()),
})
.await
});
let drop_store = store.clone();
let drop_table = tokio::spawn(async move { drop_store.drop_table(bucket, "sales", "orders").await });
let create_store = store.clone();
let mut replacement = test_table_entry(
bucket,
&source_namespace,
&source_table,
default_table_metadata_file_path(&source_namespace, &source_table, "00002.metadata.json"),
);
replacement.table_id = "replacement-table-id".to_string();
replacement.table_uuid = "replacement-table-uuid".to_string();
replacement.warehouse_location = "s3://analytics/tables/replacement-table-id".to_string();
let create = tokio::spawn(async move { create_store.create_table(replacement).await });
tokio::time::timeout(TABLE_CATALOG_TEST_TIMEOUT, async {
while backend.write_lock_acquisition_count(bucket, &publication_lock).await < publication_attempts + 3 {
tokio::task::yield_now().await;
}
})
.await
.expect("concurrent mutations should reach the table-bucket publication fence");
assert!(!commit.is_finished());
assert!(!drop_table.is_finished());
assert!(!create.is_finished());
commit.abort();
drop_table.abort();
create.abort();
rename.abort();
source_tombstone_pause.release();
let _ = rename.await;
let destination_object = store.paths.table_entry_path(
bucket,
&destination_namespace,
&IdentifierSegment::parse("orders_v2").expect("destination table should parse"),
);
let source_fence = store
.read_entry::<TableEntry>(RUSTFS_META_BUCKET, &source_object)
.await
.unwrap()
.expect("source rename fence should be durable")
.0;
let destination_fence = store
.read_entry::<TableEntry>(RUSTFS_META_BUCKET, &destination_object)
.await
.unwrap()
.expect("destination rename fence should be durable")
.0;
assert_eq!(source_fence.state, TableCatalogEntryState::Renaming);
assert_eq!(destination_fence.state, TableCatalogEntryState::Renaming);
assert_matches!(
store.load_table(bucket, "curated", "orders_v2").await,
Err(TableCatalogStoreError::Unavailable(_))
);
assert_matches!(
store.rename_table(bucket, "sales", "orders", "curated", "orders_v2").await,
Err(TableCatalogStoreError::TableNotFound(_))
);
assert!(store.load_table(bucket, "sales", "orders").await.unwrap().is_none());
let destination = store
.load_table(bucket, "curated", "orders_v2")
.await
.unwrap()
.expect("recovery should finish the destination publication");
assert_eq!(destination.table_id, "table-id");
assert!(
store
.get_table_bucket(bucket)
.await
.unwrap()
.expect("table bucket should exist")
.active_rename_id
.is_none()
);
let intent_object = store.paths.table_rename_intent_path(bucket, &active_rename_id);
let intent = store
.read_entry::<TableRenameIntent>(RUSTFS_META_BUCKET, &intent_object)
.await
.unwrap()
.expect("completed rename intent should be retained as a recovery record")
.0;
assert_eq!(intent.state, TableRenameIntentState::Completed);
}
#[tokio::test]
async fn configured_object_catalog_dispatches_table_rename() {
let store =
ConfiguredTableCatalogStore::new_for_test(TestCatalogObjectBackend::default(), TableCatalogBackingMode::ObjectBacked);
@@ -17760,6 +18217,6 @@ async fn configured_object_catalog_rejects_table_rename() {
store
.rename_table("analytics", "sales", "orders", "curated", "orders_v2")
.await,
Err(TableCatalogStoreError::Unsupported(_))
Err(TableCatalogStoreError::NotFound(_))
);
}
-2
View File
@@ -221,7 +221,6 @@ test_post_object_invalid_date_format
test_post_object_invalid_request_field_value
test_post_object_missing_policy_condition
test_post_object_request_missing_policy_specified_field
test_post_object_set_key_from_filename
test_post_object_success_redirect_action
test_post_object_tags_anonymous_request
test_post_object_wrong_bucket
@@ -241,7 +240,6 @@ test_restore_object_permanent
test_restore_object_temporary
test_sse_kms_post_object_authenticated_request
test_versioned_object_acl_no_version_specified
test_versioning_multi_object_delete_with_marker_create
test_versioning_stack_delete_merkers
# Intentionally unsupported by design: ACL-related tests
+2
View File
@@ -259,6 +259,7 @@ test_versioning_bucket_multipart_upload_return_version_id
test_versioning_concurrent_multi_object_delete
test_versioning_multi_object_delete
test_versioning_multi_object_delete_with_marker
test_versioning_multi_object_delete_with_marker_create
test_versioning_obj_create_read_remove
test_versioning_obj_create_read_remove_head
test_versioning_obj_create_versions_remove_all
@@ -474,6 +475,7 @@ test_multipart_upload_resend_part
test_object_copy_canned_acl
test_object_raw_get_x_amz_expires_not_expired
test_object_raw_get_x_amz_expires_not_expired_tenant
test_post_object_set_key_from_filename
test_put_current_object_if_match
test_put_current_object_if_none_match
test_put_delete_tags
+19 -12
View File
@@ -475,25 +475,32 @@ current unsupported inventory is:
## Credential Boundary
RustFS advertises table credential scope metadata without returning reusable
storage secrets by default. `loadTable` includes the table warehouse prefix in
the response config, and the standard credentials endpoint is registered:
storage secrets by default. The standard credentials endpoint is registered:
```text
GET /v1/{prefix}/namespaces/{namespace}/tables/{table}/credentials
```
The endpoint returns an empty `storage-credentials` list unless table catalog
credential vending is explicitly enabled. When enabled, RustFS issues temporary
table-scoped S3 credentials through the credentials endpoint. Those credentials
are constrained to the table warehouse prefix and include a session token and
expiration.
credential vending is explicitly enabled. LoadTable uses the same issuer when
the request includes `X-Iceberg-Access-Delegation: vended-credentials`. The
response advertises the issued session for the table warehouse prefix and for
the exact current metadata object; the session policy keeps table data access
inside the warehouse and grants only `GetObject` to that metadata object.
The `rustfs-vended-credentials` profile verifies the client handoff from the
catalog principal to the table-scoped temporary credentials. It still uses the
configured principal for setup and REST request signing; the vended credentials
are first checked against the created table warehouse location, then checked
with a direct S3 scope probe, and finally applied to PyIceberg S3 data-plane
access after the table has been created.
LoadTable remains metadata-only when delegation is absent, vending is disabled,
or the caller lacks the separate table-credentials permission. Disabled and
not-authorized fallbacks include an explicit reason. Issuer errors, including
disallowed chained temporary credentials, are returned as request errors rather
than silently falling back.
The `rustfs-vended-credentials` profile verifies the client handoff through the
dedicated credentials endpoint. It still uses the configured principal for
setup and REST request signing; the vended credentials are checked against the
created table warehouse location, probed directly against S3 scope boundaries,
and then applied to PyIceberg data-plane access. Stable PyIceberg releases up to
0.11 do not consume LoadTable `storage-credentials`; native LoadTable coverage
requires a client release with that support.
Enablement is server-side and fail-closed: