Skip to content

Identify resources consistently, ideally the same way gsutil does #62

Description

@jboynes

The StorageService API identifies objects using a either (bucket, name) combination e.g. the reader(bucket, name) method or as using a Blob object e.g. the writer(Blob) method.

We should use a consistent representation of an object's identity. It would be nice if that also have a string representation that was consistent with how objects are identified in other tools such as gsutil.

We could use a URI for this:

  URI id = URI.create("gs://bucket/path/to/object");
  storage.reader(id);
  storage.writer(id);

Activity

  1. aozarov commented on May 14, 2015

    @aozarov
    Contributor

    writers accept Blob because it needs (and write/update) the data in addition to the name.
    Similar to Datastore get(key), put(Entity). Blob already contains the name.

  2. modified the milestone: on Jun 1, 2015
  3. aozarov commented on Jun 1, 2015

    @aozarov
    Contributor

    As commented above, write needs the Blob. I think the place of URI for reference is in higher level libraries (such the one based on java.nio.Files API which we plan to provide in gcloud-java-contrib)

  4. jgeewax commented on Jun 3, 2015

    @jgeewax

    I'd agree that a URI might not be the best way to handle this, however I don't think we should close this out yet (until we get some responses). Reopening for debate.

  5. reopened this on Jun 3, 2015
  6. jboynes commented on Jun 3, 2015

    @jboynes
    Author

    The issue here is that we identify things in different ways across the API surface and we should make that consistent. Examples in the Storage interface:

    BlobInfo create(BlobInfo blobInfo, byte[] content, BlobTargetOption... options)
    BlobInfo update(BlobInfo blobInfo, BlobTargetOption... options)
    boolean delete(String bucket, String blob, BlobSourceOption... options)
    
    BlobInfo copy(CopyRequest copyRequest)
    
    BlobReadChannel reader(String bucket, String blob, BlobSourceOption... options)
    BlobWriteChannel writer(BlobInfo blobInfo, BlobTargetOption... options)

    I'm suggesting we clearly separate to concept of identity from data and metadata. There is a well defined way to identify an object: the structured string passed to gsutil containing the bucket and path tokens (where path can include the version). URI is simply a standard way to handle such structured strings; we could use a specific class for that e.g. ObjectId(bucket, path) but I don't see what it adds and think we'd just end up with something very close to URI. I would prefer that though to the option of passing two strings around. URI also supports relative references so could be used consistently at a global level where it would be absolute and include the bucket name, or at a per-bucket level where it would just contain the relative path component.

    BlobInfo would then become a value object representing an object's metadata independent of its identity. This simplifies server-side operations like copy or delete where the client may not be interested in the metadata at all.

    Content operations may or may not need access to metadata. For example, a download operation might want it so that it can set headers on an HTTP Response (e.g. content-length, etag), whereas an application loading/storing data for its own use might not. However, every operation would need the identity.

  7. jgeewax commented on Jun 3, 2015

    @jgeewax

    I'm suggesting we clearly separate to concept of identity from data and metadata.

    I agree with this premise -- Identity, Metadata, and Data are three separate things.

    we could use a specific class for that e.g. ObjectId(bucket, path) but I don't see what it adds and think we'd just end up with something very close to URI

    I actually would prefer the specific class for that, as Amazon does with AmazonS3URI which comes with special S3-specific methods (getKey(), getBucket()).

    I'd also argue that we make this easier to construct to avoid people manually crafting the "gs://...." string:

    // These should be equivalent.
    StorageURI uri = StorageURI("bucket", "path/to/file.txt");
    StorageURI uri = StorageURI("gs://bucket/path/to/file.txt");

    And it'd be awesome to include the helper methods...

    BucketInfo bucket = uri.getBucketInfo();
    String bucketName = uri.getBucketName();
    ObjectInfo object = uri.getObjectInfo();
    String objectName = uri.getObjectName();
  8. aozarov commented on Jun 3, 2015

    @aozarov
    Contributor

    We use BlobInfo when we need both the name and its metadata.
    We use just name when we don't need the metadata.

    I don't think the use of URI is that popular, even in cases like this (specific reference. name or bucket/name).

    Yes, Amazon do have AmazonS3URI but...

    • It is not a real URI but rather their own class that has a mong other identify method getURI
    • Most APIs that I looked at, on AmazonS3 or its requests (e.g. CopyObjectRequest, CreateBucketRequest,..) don't use the URI as input but rather bucket name and/or object name.

    I think the reason Amazon API is doing it this way (input does not seem to use URI) is for the same
    reason I am not use it (it is less convenient). However, it is totally fine to have a way to get URI from
    Bucket or Blob and to have a way to get them back from a URI. If we choose this route I think we need
    either to update the issue description or create a separate issue for it.

  9. jboynes commented on Jun 3, 2015

    @jboynes
    Author

    @jgeewax OK, sold. Not sure about StorageURI as a name - I'd suggest ObjectId ObjectName or BlobId as alternatives.

    BucketInfo bucket = uri.getBucketInfo()

    means that the id implementation has to be able to access the service to obtain the info so is no longer a pure, immutable value object. I'd suggest we don't do that and instead do something more like

    // pick a bucket
    Bucket bucket = storage.getBucket("bucket");
    BucketInfo bucketInfo = bucket.getInfo();
    
    // get metadata about an object in the bucket 
    BlobInfo info = bucket.getInfo(ObjectId.create("object"));
    
    // or as a single operation
    BlobInfo info = storage.getInfo(ObjectId.create("bucket", "object"));
  10. jgeewax commented on Jun 3, 2015

    @jgeewax

    What I'm worried about is .... "What if I want a way to pass around a unique identifier (which is effectively a bucket name and a object name), but don't want the giant list of permissions and other metadata that comes with it?"

    It seems like it might make sense to have all those methods accept...

    • the StorageURI type (which is the combo-meal of bucket name and object name),
    • the BucketInfo/ObjectInfo instances, or
    • the two String arguments as we have now.

    These might all just be shortcuts that all turn into the StorageURI class (as it's job is "identity"), but I agree that we shouldn't require people to jump through hoops and create one of these just to call delete() or reader().

    So this would mean we'd have:

    boolean delete(String bucket, String blob, BlobSourceOption... options)
    boolean delete(BlobInfo blobInfo, BlobSourceOption... options)
    boolean delete(StorageURI blobURI, BlobSourceOption... options)

    where the first two would effectively do:

    boolean delete(String bucket, String blob, BlobSourceOption... options) {
      return delete(new StorageURI(bucket, blob), options);
    }
    
    boolean delete(BlobInfo blobInfo, BlobSourceOption... options) {
      return delete(blobInfo.getStorageURI(), options);
    }
  11. jgeewax commented on Jun 3, 2015

    @jgeewax

    What about ObjectURI? I do like the idea of a URI as the primary job is to act as an identifier -- and people likely understand what it is...

  12. mziccard commented on Oct 11, 2015

    @mziccard
    Contributor

    Let me restart the discussion so that we can finally close this issue.
    I have been using Amazon's S3 API and I never needed/wanted to use AmazonS3URI.
    Do we really need an ObjectURI? While I understand the point of creating a semantic way of representing an object identify separated from it's metadata, what is the value added by doing this through a URI?

    What I propose is having a BlobId object which wraps bucketName and blobName as strings. BlobId is immutable and does not provide functional methods (no getInfo()). Rather we modify methods in Storage to accept a BlobId object:

    get(BlobId blobId, BlobSourceOption... options)
    boolean delete(BlobId blobId, BlobSourceOption... options)
    BlobReadChannel reader(BlobId blobId, BlobSourceOption... options)

    As @jgeewax suggested I would not remove current methods but rather implement them using the above ones.
    All other methods that use blob's metadata keep having a BlobInfo parameter. If you prefer we can overload those methods to take a BlobId and skip metadata checks.

    Thoughts?

  13. 16 remaining items

  14. added a commit that references this issue on Dec 22, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

🚨 criticalP0 critical issue. Requires immediate fixapi: storageIssues related to the Cloud Storage API.triage meI really want to be triaged.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions