Skip to content

docs: clarify GetObject stream ContentLength is not plaintext length - #211

Open
reginaldalfret wants to merge 1 commit into
aws:mainfrom
reginaldalfret:docs-164-stream-content-length
Open

reginaldalfret wants to merge 1 commit into
aws:mainfrom
reginaldalfret:docs-164-stream-content-length

Conversation

@reginaldalfret

Copy link
Copy Markdown

Fixes #164

Summary

Clarifies in documentation and docstrings that the ContentLength in the response returned by get_object reflects the length of the ciphertext stream stored in S3 (which includes the cryptographic authentication tag, e.g. 16 bytes for AES-GCM, or cipher padding for CBC mode), rather than the decrypted plaintext length.

Customers and callers should always read the entire stream (response["Body"].read()) to ensure complete decryption and cryptographic authentication verification rather than relying on ContentLength as the plaintext length.

Changes

  1. README.md: Added a note under Getting Started detailing the distinction between stream length (ContentLength) and decrypted plaintext length.
  2. docs/index.rst: Added a Sphinx note block covering stream length vs plaintext length and the requirement to read the full stream.
  3. src/s3_encryption/__init__.py:
    • Added a Note: section to S3EncryptionClient.get_object docstring explaining the ContentLength semantics.
    • Updated on_get_object_after_call docstring.
  4. src/s3_encryption/pipelines.py: Updated GetEncryptedObjectPipeline.decrypt docstring to document that the surrounding response's ContentLength is the ciphertext stream length.

Verification

  • ruff check src passed clean (including docstring style rules).
  • ruff format --check src passed clean.
  • Unit test suite (pytest test --ignore test/integration --ignore test/performance): 324 passed.
  • git diff --check passed clean with zero whitespace issues.

Fixes aws#164

Clarify in the README, Sphinx documentation (docs/index.rst), and API docstrings (S3EncryptionClient.get_object, on_get_object_after_call, and GetEncryptedObjectPipeline.decrypt) that the ContentLength in the response dictionary returned by get_object corresponds to the ciphertext stream length stored in S3 (which includes the cryptographic authentication tag or CBC padding), rather than the decrypted plaintext length.

Callers must read the entire stream (e.g. via response['Body'].read()) to ensure complete decryption and authentication verification.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Make sure that the Docs clarify stream length is not always plaintext length

1 participant