Multipart upload: how large files get to object storage reliably
Kort antwoord
Multipart upload splits a large file into separate chunks that upload independently and get reassembled into the final object once every part has arrived. If one chunk fails partway through, only that chunk needs retrying, not the entire file. Most S3-compatible tools and SDKs switch to multipart automatically above a certain file size, so in normal use you don't need to trigger it yourself.
How it works
A single, straightforward upload sends a file to object storage as one continuous stream from start to finish. That works fine for small files, but it has an obvious weakness for large ones: if the connection drops at 95 percent, the whole upload has failed, and the usual recovery is to start again from byte zero.
Multipart upload avoids that by breaking the file into a series of parts before it ever leaves your machine. Each part uploads on its own, as its own request, and the object storage service holds onto the parts it receives without treating the object as complete. Once every part has arrived, you (or your client tool) tell the service the upload is finished, and it assembles the parts, in order, into a single object. From the outside, once that assembly step completes, the result is indistinguishable from an object that was uploaded in one go.
Why this matters for reliability
The practical benefit is entirely about failure handling. Networks drop connections, especially over long-running transfers, and a dropped connection partway through a large single-stream upload normally means the whole transfer is wasted, with nothing to show for the time and bandwidth already spent. Multipart upload changes the unit of failure from "the whole file" to "one part." If a chunk fails to upload, only that chunk needs retrying. Everything that already succeeded stays in place, waiting for the rest to catch up, rather than being discarded along with it.
This also opens the door to uploading parts in parallel rather than strictly one after another, which is one reason multipart upload tends to be noticeably faster for very large files on top of being more resilient.
When it kicks in
You rarely need to think about any of this directly. Most S3-compatible client tools and SDKs, the AWS CLI, rclone, and boto3 among them, decide automatically whether a given upload should go as a single request or be split into parts, based on the file's size crossing an internal threshold built into the tool. Below that threshold, you get a normal single upload. Above it, the tool handles the splitting, the parallel part uploads, and the final assembly step itself, without changing anything about how you invoke it. Multipart upload is worth understanding conceptually mainly so a failed large upload, and its automatic retry of just one part rather than the whole file, doesn't come as a surprise.