EchoTwin Privacy-Aware Personalized Email Reply Generation with Low Rank Adaptation

Authors

  • Thomas Samuel Nahda University image/svg+xml Author
    Competing Interests

    NA

  • Mohammed Abdelwahab Nahda University image/svg+xml Author
    Competing Interests

    NO

  • Hanaa Atwa Nahda University image/svg+xml Author
    Competing Interests

    NO

  • Anasimon Samir Nahda University image/svg+xml Author
    Competing Interests

    NO

  • Carol Faried Nahda University image/svg+xml Author
    Competing Interests

    NO

DOI:

https://doi.org/10.66279/g71jd323

Keywords:

EchoTwin, Privacy-Aware Architecture, Personalized Email-Reply, Low-Rank Adaptation, Large Language Models

Abstract

Email continues to serve as a primary medium for both professional and personal communication. Although contemporary assistance systems, including Smart Reply and Smart Compose, enhance writing efficiency, they generally do not explicitly model or adapt to an individual’s unique communication patterns. This paper presents EchoTwin, an AI-assisted email-reply generation system that produces personalized draft replies aligned with a user’s historical communication patterns. EchoTwin builds on a decoder-only Transformer large language model and applies Low-Rank Adaptation (LoRA) for parameter-efficient personalization: a lightweight adapter is trained for each user while a single base model is shared across the population, avoiding the cost of an independent model per user. The system integrates with Gmail through Google OAuth 2.0 and the Gmail API, keeping a human reviewer in the loop to approve, edit, or reject every draft before it is sent. The approach was evaluated on a curated subset of the Enron Email-Reply dataset, comprising 2,221 email-reply pairs from ten users, comparing personalized generation against the base model on 223 held-out samples using lexical similarity (ROUGE), semantic similarity, corpus-level style similarity, and stylometric analysis. The personalized adapters achieved statistically significant gains of 8.10% in ROUGE-1, 35.65% in ROUGE-L, and 5.64% in corpus-level style similarity, plus a 25.26% reduction in mean stylometric distance (Wilcoxon signed-rank test, ???? < 0.05). Sentence-BERT similarity to the reference reply decreased by 10.29%, reflecting the expected trade-off between reference-based similarity and individual style adaptation. These prototype-level results indicate that user-specific LoRA adapters can improve measured stylistic alignment in personalized email-reply generation while keeping the user in control of the final output.

Downloads

Download data is not yet available.

Author Biographies

  • Thomas Samuel, Nahda University

    Faculty of Computer Science, Nahda University, Beni-Suef City, 62511, Egypt;

  • Mohammed Abdelwahab, Nahda University

    Faculty of Computer Science, Nahda University, Beni-Suef City, 62511, Egypt;

  • Hanaa Atwa, Nahda University

    Faculty of Computer Science, Nahda University, Beni-Suef City, 62511, Egypt.

  • Anasimon Samir, Nahda University

    Faculty of Computer Science, Nahda University, Beni-Suef City, 62511, Egypt.

  • Carol Faried, Nahda University

    Faculty of Computer Science, Nahda University, Beni-Suef City, 62511, Egypt.

References

[1] McKinsey Global Institute, “The social economy: Unlocking value and productivity through social technologies,” tech. rep., McKinsey & Company, New York, NY, USA, 2012.

[2] M. Plummer, “How to spend way less time on email every day,” Harvard Business Review, vol. 22, 2019.

[3] A. Kannan, K. Kurach, S. Ravi, T. Kaufmann, A. Tomkins, B. Miklos, G. Corrado, L. Lukacs, M. Ganea, P. Young, et al., “Smart reply: Automated response suggestion for email,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pp. 955–964, 2016. DOI: https://doi.org/10.1145/2939672.2939801

[4] M. X. Chen, B. N. Lee, G. Bansal, Y. Cao, S. Zhang, J. Lu, J. Tsay, Y. Wang, A. M. Dai, Z. Chen, et al., “Gmail smart compose: Real-time assisted writing,” in Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 2287–2295, 2019. DOI: https://doi.org/10.1145/3292500.3330723

[5] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025.

[6] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685, 2021.

[7] Google Developers, “Using OAuth 2.0 to access Google APIs.” https://developers.google.com/identity/ protocols/oauth2, 2026. Accessed: May 11, 2026.

[8] Google Developers, “Gmail API Reference: users.messages.send.” Google for Developers. [Online]. Available: https://developers.google.com/gmail/api/reference/rest, 2026. Accessed: May 11, 2026.

[9] OWASP Foundation, “OWASP Top Ten Web Application Security Risks.” OWASP. [Online]. Available: https://owasp.org/www-project-top-ten/, 2025. Accessed: May 12, 2026.

[10] E. Tabassi, “Artificial intelligence risk management framework (ai rmf 1.0),” Tech. Rep. NIST AI 100-1, National Institute of Standards and Technology, Gaithersburg, MD, USA, 2023. DOI: https://doi.org/10.6028/NIST.AI.100-1

[11] A. Silva, P. Tambwekar, and M. Gombolay, “Fedperc: Federated learning for language generation with personal and context preference embeddings,” in Findings of the Association for Computational Linguistics: EACL 2023, pp. 869–882, 2023. DOI: https://doi.org/10.18653/v1/2023.findings-eacl.64

[12] A. Salemi, S. Kallumadi, and H. Zamani, “Optimization methods for personalizing large language models through retrieval augmentation,” in Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 752–762, 2024. DOI: https://doi.org/10.1145/3626772.3657783

[13] A. Salemi and H. Zamani, “Comparing retrieval-augmentation and parameter-efficient fine-tuning for privacy-preserving personalization of large language models,” in Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR), pp. 286–296, 2025. DOI: https://doi.org/10.1145/3731120.3744595

[14] Z. Tan, Q. Zeng, Y. Tian, Z. Liu, B. Yin, and M. Jiang, “Democratizing large language models via personalized parameter-efficient fine-tuning,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 6476–6491, 2024. DOI: https://doi.org/10.18653/v1/2024.emnlp-main.372

[15] C. Sun, K. Yang, R. G. Reddy, Y. Fung, H. P. Chan, K. Small, C. Zhai, and H. Ji, “Persona-db: Efficient large language model personalization for response prediction with collaborative data refinement,” in Proceedings of the 31st International Conference on Computational Linguistics, pp. 281–296, 2025.

[16] Y. Xu, J. Zhang, A. Salemi, X. Hu, W. Wang, F. Feng, H. Zamani, X. He, and T.-S. Chua, “Personalized generation in large model era: A survey,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 24607–24649, 2025. DOI: https://doi.org/10.18653/v1/2025.acl-long.1201

[17] Y. Li, Z. Tan, P. Branco, and Y. Liu, “Privacy-preserving parameter-efficient fine-tuning for large language model services,” IEEE Transactions on Audio, Speech and Language Processing, 2025. DOI: https://doi.org/10.1109/TASLPRO.2025.3612842

[18] Y. Wang, Y. Lin, X. Zeng, and G. Zhang, “Privatelora for efficient privacy preserving llm,” arXiv preprint arXiv:2311.14030, 2023.

[19] Z. Zeng, J. Wang, J. Yang, Z. Lu, H. Li, H. Zhuang, and C. Chen, “Privacyrestore: Privacy-preserving inference in large language models via privacy removal and restoration,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 10821–10855, 2025. DOI: https://doi.org/10.18653/v1/2025.acl-long.532

[20] R. Singhal, K. Ponkshe, and P. Vepakomma, “Fedex-lora: Exact aggregation for federated and efficient fine-tuning of large language models,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1316–1336, 2025. DOI: https://doi.org/10.18653/v1/2025.acl-long.67

[21] Z. Shen, J. Lu, H. Wan, and J. Chen, “Sdflora: Selective decoupled federated lora for privacy-preserving fine-tuning with heterogeneous clients,” arXiv preprint arXiv:2601.11219, 2026.

[22] J. Liu, Y. Miao, N. Xi, and J. Liu, “Rethinking lora for privacy-preserving federated learning in large models,” arXiv preprint arXiv:2602.19926, 2026.

[23] A. Salemi, S. Mysore, M. Bendersky, and H. Zamani, “Lamp: When large language models meet personalization,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 7370–7392, 2024. DOI: https://doi.org/10.18653/v1/2024.acl-long.399

[24] I. Kumar, S. Viswanathan, S. Yerra, A. Salemi, R. A. Rossi, F. Dernoncourt, H. Deilamsalehy, X. Chen, R. Zhang, S. Agarwal, et al., “Longlamp: A benchmark for personalized long-form text generation,” arXiv preprint arXiv:2407.11016, 2024.

[25] B. Jiang, Z. Hao, Y.-M. Cho, B. Li, Y. Yuan, S. Chen, L. Ungar, C. J. Taylor, and D. Roth, “Know me, respond to me: Benchmarking llms for dynamic user profiling and personalized responses at scale,” arXiv preprint arXiv:2504.14225, 2025.

[26] X. Fu, H. A. Rahmani, B. Wu, J. Ramos, E. Yilmaz, and A. Lipani, “Pref: Reference-free evaluation of personalised text generation in llms,” arXiv preprint arXiv:2508.10028, 2025.

[27] J. Park, D. Kim, and T. Moon, “Prisp: Privacy-safe few-shot personalization via lightweight adaptation,” in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 24986–25003, 2026. DOI: https://doi.org/10.18653/v1/2026.acl-long.1146

[28] X. Li, R. Zhou, Z. C. Lipton, and L. Leqi, “Personalized language modeling from personalized human feedback,” arXiv preprint arXiv:2402.05133, 2024.

[29] Kaggle, “Enron email-reply dataset.” Kaggle. [Online]. Available: https://www.kaggle.com/datasets/ oanannv/enron-email-reply-dataset, 2023. Accessed: May 17, 2026.

[30] Carnegie Mellon University, “Enron email dataset.” CMU School of Computer Science. [Online]. Available: https://www.cs.cmu.edu/~enron/, 2015. Accessed: May 17, 2026.

[31] N. Sakimura, J. Bradley, and N. Agarwal, “Proof key for code exchange by oauth public clients,” Tech. Rep. RFC 7636, RFC Editor, Sept. 2015. DOI: https://doi.org/10.17487/RFC7636

[32] N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pp. 3982–3992, 2019. DOI: https://doi.org/10.18653/v1/D19-1410

[33] P. Juola, “Authorship attribution,” Foundations and Trends® in Information Retrieval, vol. 1, no. 3, pp. 233–334, 2008. DOI: https://doi.org/10.1561/1500000005

[34] M. Koppel, J. Schler, and S. Argamon, “Authorship attribution in the wild,” Language Resources and Evaluation, vol. 45, no. 1, pp. 83–94, 2011. DOI: https://doi.org/10.1007/s10579-009-9111-2

[35] React Documentation, “React documentation.” Online, 2026. Available: https://react.dev/learn. Accessed: May 17, 2026.

[36] S. Tiangolo, “Fastapi documentation,” URL: https://fastapi. tiangolo. com, 2025.

[37] B. Ward, SQL Server 2025 Unveiled: The AI-Ready Enterprise Database with Microsoft Fabric Integration. Berkeley, CA, USA: Apress, 2025. DOI: https://doi.org/10.1007/979-8-8688-1847-9

[38] H. Face, “Transformers documentation,” URL: https://huggingface. co/docs/transformers/index, 2024.

[39] C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out, pp. 74–81, 2004.

[40] F. Wilcoxon, “Individual comparisons by ranking methods,” Biometrics bulletin, vol. 1, no. 6, pp. 80–83, 1945. DOI: https://doi.org/10.2307/3001968

Downloads

Published

29-08-2026

Data Availability Statement

The data underlying this study are drawn from the Enron Email-Reply dataset, a curated subset of the original Enron email corpus, both of which are publicly available: the Enron Email-Reply dataset is distributed on Kaggle [29 ], and the original Enron corpus is distributed by Carnegie Mellon University [30].

How to Cite

EchoTwin Privacy-Aware Personalized Email Reply Generation with Low Rank Adaptation. (2026). Computational Discovery and Intelligent Systems (CDIS), 5(1), 1-25. https://doi.org/10.66279/g71jd323

Similar Articles

1-10 of 19

You may also start an advanced similarity search for this article.