Skip to main content

Design a Cloud File Storage System

Design a cloud file storage system with file sync, deduplication, sharing, versioning, and conflict resolution across devices.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a distributed systems architect. Design a cloud file storage system similar to Dropbox supporting 100 million users with 2 TB storage each.

Step 1: Define the core features: file upload and download, automatic sync across 5 devices per user, file and folder sharing with permission levels (view, edit, admin), file versioning with 180 days version history, and real-time collaborative editing. Non-functional requirements: upload speed of 100 MB/s for large files, sync latency under 3 seconds for small file changes, and 99.999999999% data durability.

Step 2: Design the file storage architecture. Split files into 4 MB chunks for upload, storage, and sync. Store chunks in Amazon S3 with 3x replication across availability zones. Implement content-addressable storage using SHA-256 hashes for chunk-level deduplication across all users. Calculate the expected deduplication ratio and storage savings for 100 million users.

Step 3: Design the metadata service that tracks the file hierarchy, permissions, and sync state. Store metadata in MySQL with sharding with a schema supporting: file/folder tree per user, sharing relationships, version chains, and per-device sync cursors. The sync cursor tracks which changes each device has received. Design the schema to support listing folder contents with under 100ms latency.

Step 4: Implement the sync protocol. Each client maintains a local database of file metadata and sync state. On file change, the client uploads new/modified chunks and sends a metadata update to the server. The server increments a global version counter and notifies all other devices via long polling. Each device pulls changes since its last sync cursor. Handle offline edits by queueing changes and replaying on reconnect.

Step 5: Design conflict resolution for when two devices edit the same file while offline. Detect conflicts by comparing parent version IDs. For document files, attempt automatic merge using operational transformation. For binary files, save both versions as conflict copies with clear naming. Let the user choose which version to keep via the desktop app dialog interface.

Step 6: Build the sharing and collaboration layer. Implement link-based sharing with configurable permissions and expiration dates. Design the permission model to support nested folder permissions with inheritance. For real-time collaboration, integrate an operational transformation or CRDT engine that enables 50 simultaneous editors on a single document.

What this prompt does

This prompt makes the AI a distributed systems architect designing a cloud file storage system similar to [reference_platform] for [user_count] users with [storage_per_user] each. It works through storage architecture, a metadata service, the sync protocol, conflict resolution, and a sharing and collaboration layer. Files are split into [chunk_size] chunks stored in [object_storage] with [replication_factor] replication, and content-addressable hashing enables chunk-level deduplication across all users.

The structure works because file sync breaks in predictable places: chunking, deduplication, and conflict handling. Forcing the deduplication math up front matters because the expected savings change the storage backend decision more than people expect, and chunk-level hashing means identical content stored by different users only costs space once. The sync protocol uses per-device cursors so each device pulls only changes since its last sync rather than re-downloading the tree, and offline edits queue locally and replay on reconnect. Conflict resolution compares parent version IDs — attempting operational-transform merges for documents and saving clearly named conflict copies for binaries via [conflict_ui] so nothing is silently lost.

When to use it

  • You're designing a sync or storage layer and want chunking and dedup decided early
  • You need a metadata model that tracks the file tree, permissions, and per-device sync state
  • You want a sync protocol that handles offline edits and replays on reconnect
  • You're tackling conflict resolution for simultaneous offline edits
  • You need sharing with nested-folder permission inheritance and link expiration
  • You're adding real-time collaboration with multiple simultaneous editors
  • You want a deduplication and durability estimate at [user_count] scale before choosing a backend

Example output

Expect a layered design: a storage section with [chunk_size] chunking, content-addressable dedup via SHA-256, [replication_factor] replication across availability zones, and a savings estimate; a metadata schema in [metadata_database] covering the file tree, sharing relationships, version chains, and per-device sync cursors; the client-server sync protocol with cursor-based change pulls and offline replay; a conflict-resolution flow distinguishing document merges from binary conflict copies; and a sharing layer with link permissions, expiration dates, nested inheritance, and a CRDT or OT engine for up to [max_collaborators] simultaneous editors. It's an architecture walkthrough, not code.

Pro tips

  • Run the deduplication math early; the realistic dedup ratio at [user_count] scale can reshape your [object_storage] choice
  • Pick [chunk_size] to balance dedup granularity against metadata overhead — smaller chunks dedup better but multiply bookkeeping
  • Lean on per-device sync cursors so devices pull deltas, not full trees, on every change
  • Distinguish document and binary conflict handling clearly; OT merges suit text, conflict copies suit binaries
  • Set [replication_factor] to meet your durability goal across availability zones, not just within one
  • Design nested-permission inheritance carefully; ambiguous inheritance rules are a common source of sharing bugs

Frequently Asked Questions

How does deduplication save storage across users?
Files are split into `[chunk_size]` chunks and each chunk is identified by a SHA-256 hash. When two users store identical chunks, the system keeps only one copy and references it from both, so common files and shared content consume storage once rather than per user.
How does the system know which changes each device still needs?
The metadata service maintains a per-device sync cursor that records the last change each device received. On sync, a device pulls only the changes since its cursor, and offline edits are queued locally and replayed on reconnect, so devices converge without re-downloading everything.
What happens when two devices edit the same file offline?
Conflicts are detected by comparing parent version IDs. For document files the system attempts an automatic merge using operational transformation, and for binary files it saves both versions as clearly named conflict copies so nothing is silently lost. The user resolves the choice through the `[conflict_ui]` interface.
Does this support real-time collaborative editing?
Yes. The sharing and collaboration layer integrates an operational-transformation or CRDT engine that supports up to `[max_collaborators]` simultaneous editors on a single document. That's separate from the offline conflict resolution, which handles the case where edits happen while disconnected.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in System Design Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support