To address the issue in the existing Whole Slide Image (WSI) classification for lung cancer, where Multi-Instance Learning (MIL) methods based on knowledge distillation architecture struggle to handle the mixture of hard and easy instances effectively, leading to insufficient model learning capability and imbalanced performance, a lung cancer whole slide image classification method based on Bidirectional Collaborative Distillation and Multi-Instance Learning (BCD-MIL) was proposed by incorporating the concept of multi-task learning. First, a Bidirectional Collaborative Distillation Framework (BCDF) comprising multiple student models and a teacher model was designed, so as to satisfy the differentiated learning requirements for easy and hard pathological instances in lung cancer histopathological images, and by distributing different tasks to multiple student models, the model's learning capacity was enhanced while solving performance imbalance problem. Concurrently, classification performance was improved through a virtuous cycle with the teacher model. Second, a Multi-Angle Instance Mining Parallel architecture (MA-IMP) was designed to match the heterogeneous differences in lung cancer pathological features across dimensions such as cell morphology and tissue texture, and by conducting instance mining from multiple perspectives, the mining bias caused by a single perspective was avoided. Finally, a Dynamic Stage-Aware Exponential Moving Average distillation (DSA-EMA) algorithm was proposed to optimize weight update of the teacher model, and improve training efficiency and model performance based on the stage characteristics of large-scale instance training of lung cancer histopathological images through adjusting distillation parameters in the training phase dynamically. Experimental results show that compared to the MIL framework with Masked Hard Instance Mining (MHIM-MIL) method, which is also based on knowledge distillation and instance mining, BCD-MIL achieves improvements of 1.92, 2.37, 3.43, and 1.75 percentage points in Area Under Curve (AUC), accuracy, F1-Score, and recall, respectively, on The Cancer Genome Atlas (TCGA) dataset, and improvements of 1.11, 6.29, 6.94, and 12.85 percentage points in four metrics, respectively, on the Clinical Proteomic Tumor Analysis Consortium (CPTAC) dataset; validating the effectiveness of the proposed method. Furthermore, the lightweight distillation architecture reduces model parameter size and inference time, thereby enabling efficient deployment while ensuring performance gains, and providing a reliable basis for lung cancer WSI classification.