An architecture comparison points to a trade-off between learning new tasks and retaining earlier ones: width and global pooling may affect retention, while some networks learn quickly but forget more.
The source summarizes a comparison of architectures for continual learning. According to the publication, wider networks and removing or reducing global average pooling can improve retention of previously learned knowledge.
ResNets and WideResNets, by contrast, are described as quick to learn new tasks but more prone to forgetting earlier learning. Architecture choices can therefore affect the balance between stability and plasticity, in addition to the learning algorithm.
To assess these options, define separate metrics for performance on the new task and retention on previous ones. Compare architectures under the same conditions and track results across tasks rather than drawing conclusions from a single test.
If you use AI to study or apply the material, avoid submitting unnecessary personal data or internal information. The reported finding is a result to investigate, not a guarantee that an architecture will have the same effect on another set of tasks.