Skip to content

Support style-transfer models with changes/additions to the following operations - #123

Merged
wchao1115 merged 6 commits into
masterfrom
wchao/style_transfer
Dec 22, 2020
Merged

wchao1115 merged 6 commits into
masterfrom
wchao/style_transfer

Conversation

@wchao1115

@wchao1115 wchao1115 commented Nov 28, 2020 •

Copy link
Copy Markdown
Collaborator

#108

  • Extend the conv2d operation to support transposed convolution, an essential upsample tool for encoder-decoder models.
  • Add instanceNormalization in addition to batch-normalization. Instance-normalization is a fused operation for a normalization subgraph that computes the mean and variance values per-feature instance on the fly.
  • Replace the sqrt unary operation with a more generic pow binary operation used by the normalization process.
  • Add pad operation that supports all 4 padding modes found in various frameworks.
  • Add resample operation to support both upsampling and downsampling of feature instances. This operation is used in the ONNX version of the style-transfer models.

Preview | Diff

- Extend the `conv2d` operation to support transposed convolution, an essential upsample tool for encoder-decoder models.
- Add `instanceNormalization` in addition to batch-normalization. Instance-normalization is a fused operation for a normalization subgraph that computes the mean and variance values per-feature instance on the fly.
- Replace the `sqrt` unary operation with a more generic `pow` binary operation used by the normalization process.
- Add `pad` operation that supports all 4 padding modes found in various frameworks.
- Add `resample` operation to support both upsampling and downsampling of feature instances. This operation is used in the ONNX version of the style-transfer models.

@anssiko anssiko left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! Added to the 2020-12-10 agenda for discussion.

As usual, if we get all the reviews completed ahead the call we can merge this earlier than that and recap the changes on the call.

@wchao1115

Copy link
Copy Markdown
Collaborator Author

@gramalingam

Comment thread index.bs
Normalize the input features per feature instance using [[Instance-Normalization]]. Unlike [[#api-modelbuilder-batchnorm]] where the mean and variance values are computed across the batch dimension during training, the mean and variance values of instance normalization is computed from each individual feature instance on the fly.
<script type=idl>
dictionary InstanceNormalizationOptions {
Operand scale;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be better to not allow "Operands" in "Options", and make "Operands" explicit (even if optional). I also think it is useful to restrict "Options" to values that are known at the graph-building time (while "Operands" represent values known at run-time).

@wchao1115 wchao1115 Dec 1, 2020 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@gramalingam That's a good observation and I have thought about this exact point but ultimately went for the idea that any optional param, whether it's an operand or not, ought to be placed in the options struct because the original motivation for this is to simplify future changes to the API so that no overload will ever be needed in the future.

So even if we are willing to draw an arbitrary line right now barring a non-compile time param from entering the options struct, we will be forced to break that rule anyway in the future at the very first instance of the need to extend the API to support an additional optional operand.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"Any optional param ought to be placed in the options struct" seems a bit extreme. In the end, either is functionally equivalent, and it is mostly a matter of style as to which is more readable. I personally prefer making the operand explicit, but am okay with this

Comment thread index.bs
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs
…es for transposed convolution.

Add `layout` selection for instanceNormalization as both the feature and spatial dimensions need to be specified.
@wchao1115 wchao1115 changed the title Support style-transfer models with changes/additions of the following operations Support style-transfer models with changes/additions to the following operations Dec 8, 2020

@huningxin huningxin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, thanks for your great efforts @wchao1115 !

@gramalingam

Copy link
Copy Markdown
Contributor

One general question relating to padding in convolution etc.: ONNX and TF have options like SAME for padding. I don't see something similar here, is there a way to achieve that?

@wchao1115

wchao1115 commented Dec 10, 2020 •

Copy link
Copy Markdown
Collaborator Author

One general question relating to padding in convolution etc.: ONNX and TF have options like SAME for padding. I don't see something similar here, is there a way to achieve that?

Algorithmic padding can be handled by the padding array. Here, the SAME padding simply means whether the extra paddings are to be done to the left (lower end) or to the right (upper end) of the input spatial dimensions. ONNX has more flexibility in specifying which side of the input dimensions is prioritized to receive the extra padding first while TensorFlow SAME padding has a fixed pad upper logic.

That said if the input dimensions aren't known at graph building time (in models with free dimensions), then obviously the padding array can't be calculated ahead of time either, similar to the argument that @huningxin made earlier in the thread about the need for outputSizes in the transposed case. We can add an autoPad option to mitigate that issue at runtime.

…he input tensor shape isn't known at graph construction time.
@wchao1115

Copy link
Copy Markdown
Collaborator Author

@gramalingam Can you please take a look at the latest commit? Based on your feedback, an autoPad option is added to handle the case where the input shape isn't known at graph construction time.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants