Repository navigation
Identify resources consistently, ideally the same way gsutil does #62
Description
Activity
writers accept Blob because it needs (and write/update) the data in addition to the name.
Similar to Datastore get(key), put(Entity). Blob already contains the name.- addedapi: storageIssues related to the Cloud Storage API.Issues related to the Cloud Storage API.
on May 14, 2015 As commented above, write needs the Blob. I think the place of URI for reference is in higher level libraries (such the one based on java.nio.Files API which we plan to provide in gcloud-java-contrib)
I'd agree that a URI might not be the best way to handle this, however I don't think we should close this out yet (until we get some responses). Reopening for debate.
The issue here is that we identify things in different ways across the API surface and we should make that consistent. Examples in the Storage interface:
BlobInfo create(BlobInfo blobInfo, byte[] content, BlobTargetOption... options) BlobInfo update(BlobInfo blobInfo, BlobTargetOption... options) boolean delete(String bucket, String blob, BlobSourceOption... options) BlobInfo copy(CopyRequest copyRequest) BlobReadChannel reader(String bucket, String blob, BlobSourceOption... options) BlobWriteChannel writer(BlobInfo blobInfo, BlobTargetOption... options)
I'm suggesting we clearly separate to concept of identity from data and metadata. There is a well defined way to identify an object: the structured string passed to gsutil containing the bucket and path tokens (where path can include the version).
URIis simply a standard way to handle such structured strings; we could use a specific class for that e.g. ObjectId(bucket, path) but I don't see what it adds and think we'd just end up with something very close to URI. I would prefer that though to the option of passing two strings around.URIalso supports relative references so could be used consistently at a global level where it would be absolute and include the bucket name, or at a per-bucket level where it would just contain the relative path component.BlobInfo would then become a value object representing an object's metadata independent of its identity. This simplifies server-side operations like
copyordeletewhere the client may not be interested in the metadata at all.Content operations may or may not need access to metadata. For example, a download operation might want it so that it can set headers on an HTTP Response (e.g. content-length, etag), whereas an application loading/storing data for its own use might not. However, every operation would need the identity.
I'm suggesting we clearly separate to concept of identity from data and metadata.
I agree with this premise -- Identity, Metadata, and Data are three separate things.
we could use a specific class for that e.g. ObjectId(bucket, path) but I don't see what it adds and think we'd just end up with something very close to URI
I actually would prefer the specific class for that, as Amazon does with
AmazonS3URIwhich comes with special S3-specific methods (getKey(),getBucket()).I'd also argue that we make this easier to construct to avoid people manually crafting the "gs://...." string:
// These should be equivalent. StorageURI uri = StorageURI("bucket", "path/to/file.txt"); StorageURI uri = StorageURI("gs://bucket/path/to/file.txt");
And it'd be awesome to include the helper methods...
BucketInfo bucket = uri.getBucketInfo(); String bucketName = uri.getBucketName(); ObjectInfo object = uri.getObjectInfo(); String objectName = uri.getObjectName();
We use
BlobInfowhen we need both the name and its metadata.
We use just name when we don't need the metadata.I don't think the use of
URIis that popular, even in cases like this (specific reference. name or bucket/name).Yes, Amazon do have
AmazonS3URIbut...- It is not a real URI but rather their own class that has a mong other identify method getURI
- Most APIs that I looked at, on
AmazonS3or its requests (e.g.CopyObjectRequest,CreateBucketRequest,..) don't use the URI as input but rather bucket name and/or object name.
I think the reason Amazon API is doing it this way (input does not seem to use URI) is for the same
reason I am not use it (it is less convenient). However, it is totally fine to have a way to get URI from
Bucket or Blob and to have a way to get them back from a URI. If we choose this route I think we need
either to update the issue description or create a separate issue for it.@jgeewax OK, sold. Not sure about
StorageURIas a name - I'd suggestObjectIdObjectNameorBlobIdas alternatives.BucketInfo bucket = uri.getBucketInfo()means that the id implementation has to be able to access the service to obtain the info so is no longer a pure, immutable value object. I'd suggest we don't do that and instead do something more like
// pick a bucket Bucket bucket = storage.getBucket("bucket"); BucketInfo bucketInfo = bucket.getInfo(); // get metadata about an object in the bucket BlobInfo info = bucket.getInfo(ObjectId.create("object")); // or as a single operation BlobInfo info = storage.getInfo(ObjectId.create("bucket", "object"));
What I'm worried about is .... "What if I want a way to pass around a unique identifier (which is effectively a bucket name and a object name), but don't want the giant list of permissions and other metadata that comes with it?"
It seems like it might make sense to have all those methods accept...
- the
StorageURItype (which is the combo-meal of bucket name and object name), - the
BucketInfo/ObjectInfoinstances, or - the two
Stringarguments as we have now.
These might all just be shortcuts that all turn into the StorageURI class (as it's job is "identity"), but I agree that we shouldn't require people to jump through hoops and create one of these just to call
delete()orreader().So this would mean we'd have:
boolean delete(String bucket, String blob, BlobSourceOption... options) boolean delete(BlobInfo blobInfo, BlobSourceOption... options) boolean delete(StorageURI blobURI, BlobSourceOption... options)
where the first two would effectively do:
boolean delete(String bucket, String blob, BlobSourceOption... options) { return delete(new StorageURI(bucket, blob), options); } boolean delete(BlobInfo blobInfo, BlobSourceOption... options) { return delete(blobInfo.getStorageURI(), options); }
- the
What about
ObjectURI? I do like the idea of a URI as the primary job is to act as an identifier -- and people likely understand what it is...Let me restart the discussion so that we can finally close this issue.
I have been using Amazon's S3 API and I never needed/wanted to useAmazonS3URI.
Do we really need anObjectURI? While I understand the point of creating a semantic way of representing an object identify separated from it's metadata, what is the value added by doing this through aURI?What I propose is having a
BlobIdobject which wrapsbucketNameandblobNameas strings.BlobIdis immutable and does not provide functional methods (nogetInfo()). Rather we modify methods inStorageto accept aBlobIdobject:get(BlobId blobId, BlobSourceOption... options) boolean delete(BlobId blobId, BlobSourceOption... options) BlobReadChannel reader(BlobId blobId, BlobSourceOption... options)
As @jgeewax suggested I would not remove current methods but rather implement them using the above ones.
All other methods that use blob'smetadatakeep having aBlobInfoparameter. If you prefer we can overload those methods to take aBlobIdand skip metadata checks.Thoughts?
16 remaining items
- added a commit that references this issue
on Dec 22, 2025 - added 2 commits that reference this issue
on Jan 6, 2026 - added a commit that references this issue
on Feb 24, 2026 - added a commit that references this issue
on Mar 11, 2026 - added a commit that references this issue
on Mar 12, 2026 - added a commit that references this issue
on Mar 30, 2026 - added a commit that references this issue
on Apr 1, 2026
The StorageService API identifies objects using a either (bucket, name) combination e.g. the reader(bucket, name) method or as using a Blob object e.g. the writer(Blob) method.
We should use a consistent representation of an object's identity. It would be nice if that also have a string representation that was consistent with how objects are identified in other tools such as gsutil.
We could use a URI for this: