Here we conducted the experiments using the popular multi-task learning models and compare their performance.
Dataset
We used the Ali-CCP dataset for the experiments here. The details of the dataset has been introduced here.
Dataset processing
The Ali-CCP dataset would be first processed by the processing code here so that we have the artifacts having joined features for each impression and ready to be loaded for training.
- extract.py: Extract the compressed file (.tar.gz) into .csv
- data.py part 1:
parse_raw_ali_ccp_streaming_write- First read the common-features csv file and create a dictionary of mapping the common feature key to the stored common features
- Read the skeleton file by chunks size 500K rows
- For each read chunk, for each row in the chunk, compose the per impression features by joining the common features and the features stored in the skeleton
- Write the results to
parsed_{train, test}_rows_full.parquet
- data.py part 2:
load_or_biuld_sparse_vocabs_filtered_parquet,stream_normalize_parquet- For the field
featwhich is a sparse field of string, we need to build the vocabulary to keep only the high frequency value of that field and treat the rest as<UNK>and map to the vocabulary index 0- For implementation detail, we keep feat that appears at least 5 times
- The
valassociated with thefeatwill not be touched. For example, if there is a row with: field_id: 109_14, feat: “Electronic”, val: 25: the “Electronic” would be converted to<UNK>if it is not high frequency while the val 25 information would still get into the feature - Normalization: for the dense feature
val: 25part, we would applylog1p(abs(x)) * sign(x)compressing the long tail - The above is using the training dataset only, test dataset never get involved
- The output is
preprocessed_{train, test}.parquetalong withvocabwhich has 23 entries for each field id in the form likevocab["109_14"] = {"sport": 1, "electronics": 2, ...}- The
featin thepreprocessed_{train, test}.parquetis still raw string at the step
- The
- For the field
- encode.py
- This part is loaded every time the training is started. It is the doing the tensorizing
- For sparse columns, it would map the string value to an embedding index through vocabulary
- For dense value, map it to float32
- In the experiments, we have 23 sparse fields and 8 dense fields
- For label, map it to float32
Models
On the modeling side, we want to know what gains could the MTL architectures brings compared to those classic RecSys models. The experiments are designed as keeping the model architecture the only variable, other setups are the same. For the classic RecSys models, we do some tweaking to make it able to be used in the same training setup
The involved models are
- Wide&Deep
- Deep module has two heads, one is deep_ctr and the other is deep_cvr
- For the wide module, we have one wide module for ctr and another for cvr (wide_ctr, wide_cvr)
- logit_ctr = wide_ctr + deep_ctr; logit_cvr = wide_cvr + deep_cvr
- p_ctcvr = sigmoid(logit_ctr) * sigmoid(logit_cvr)
- Implementation here
- Caveat: Here the wide implementation does not use the cross operation which could have impacts on the performance. Also how to combine the deep scores and the wide scores should be more sophisticated. This is a follow up experiment
- DeepFM
- DCNv2
- Shared bottom
- MMoE
- PLE
