process_multiple_metrics.py
Introduction
This script automates the process of evaluating rewritten JSON files against multiple evaluation metrics, extracting aggregated scores, and merging the results into a single metrics file. It leverages various evaluation scripts to compute different metrics and handles the extraction and aggregation of scores.
Input/Output
Input
- Rewritten JSON Files: Files located in the
rewrittendirectory with the suffix_rewritten.json.
Output
- Metrics Files: JSON files containing individual metric scores generated by each evaluation script.
- Merged Metrics File: A JSON file that consolidates all metrics and aggregated scores for each rewritten file.
Functionality
Features
- Evaluation Script Execution: Runs a set of predefined evaluation scripts on each rewritten JSON file.
- Aggregated Score Extraction: Extracts aggregated scores from the output of each evaluation script.
- Average Aggregated Score Calculation: Calculates the average of all aggregated scores.
- Metrics Merging: Merges the individual metric scores into a single JSON file.
Code Structure
Imports
import os
import json
import glob
import subprocess
import logging
import re- os: For file and directory operations.
- json: For reading and writing JSON files.
- glob: For file pattern matching.
- subprocess: For running external evaluation scripts.
- logging: For logging information and errors.
- re: For regular expressions.
Logging Configuration
nltk_logger = logging.getLogger('nltk')
nltk_logger.setLevel(logging.ERROR)
logging.basicConfig(level=logging.DEBUG)
logger = logging.getLogger(__name__)Suppresses NLTK log messages and sets up logging for the script.
Constants
REWRITTEN_DIR = 'rewritten'
EVALUATION_SCRIPTS = [
'evaluate_cohesion_concreteness.py',
'evaluate_cohesion_information_density.py',
'evaluate_frequency.py',
'evaluate_information_density.py',
'evaluate_lexical_metrics.py',
'evaluate_lexical_syntactic_metrics.py',
'evaluate_noun_verb_metrics.py',
'evaluate_punctuation_function_words.py',
'evaluate_statistical_metrics.py',
'evaluate_structural_metrics.py'
]Defines the directory containing rewritten JSON files and the list of evaluation scripts to run.
Running Evaluation Scripts
def run_evaluation_scripts(input_file, all_aggregated_scores):
"""Run evaluation scripts on the given input file and extract aggregated scores."""
base_name = os.path.basename(input_file).replace('.json', '')
for script in EVALUATION_SCRIPTS:
logger.info("Running %s on %s", script, input_file)
result = subprocess.run(['python', script, input_file], capture_output=True, text=True)
if result.returncode != 0:
logger.error("Error running %s on %s", script, input_file)
logger.error(result.stderr)
else:
logger.info(result.stdout)
metric_file_pattern = os.path.join(REWRITTEN_DIR, f'{base_name}_metrics_{script.split("_")[1]}*.json')
metric_files = glob.glob(metric_file_pattern)
logger.debug("Pattern used for glob: %s", metric_file_pattern)
logger.info("Generated metric files: %s", metric_files)
for metric_file in metric_files:
if os.path.exists(metric_file):
logger.info("Processing metric file: %s", metric_file)
with open(metric_file, 'r') as file:
data = json.load(file)
extracted_scores = extract_aggregated_scores(data)
logger.info("Extracted aggregated scores from %s: %s", metric_file, extracted_scores)
all_aggregated_scores.extend(extracted_scores)
else:
logger.warning("Metric file %s does not exist", metric_file)Runs each evaluation script on the given input file and extracts aggregated scores from the generated metric files.
Extracting Aggregated Scores
def extract_aggregated_scores(data):
"""Extract aggregated scores from the given JSON data."""
aggregated_scores = []
if isinstance(data, dict):
for key, value in data.items():
if isinstance(value, dict):
aggregated_scores.extend(extract_aggregated_scores(value))
elif re.match(r'Aggregated.*Score', key) and isinstance(value, (int, float)):
logger.info("Found aggregated score: %s -> %s", key, value)
aggregated_scores.append((key, value))
elif isinstance(data, list):
for item in data:
aggregated_scores.extend(extract_aggregated_scores(item))
return aggregated_scoresRecursively extracts aggregated scores from the JSON data.
Calculating Average Aggregated Score
def calculate_average_aggregated_score(aggregated_scores):
"""Calculate the average of the aggregated scores."""
if aggregated_scores:
scores = [score for _, score in aggregated_scores]
logger.debug("Calculating average of aggregated scores: %s", scores)
return sum(scores) / len(scores)
logger.debug("No aggregated scores found")
return NoneCalculates the average of all extracted aggregated scores.
Merging Metrics Files
def merge_metrics_files(input_file, all_aggregated_scores):
"""Merge metrics files for the given input JSON file."""
base_name = os.path.basename(input_file).replace('.json', '')
pattern = os.path.join(REWRITTEN_DIR, f'{base_name}_metrics_*.json')
merged_metrics = {}
for file_path in glob.glob(pattern):
logger.info("Processing metrics file: %s", file_path)
with open(file_path, 'r') as file:
data = json.load(file)
logger.debug("Loaded data: %s", json.dumps(data, indent=2))
for key, value in data.items():
if key not in merged_metrics:
merged_metrics[key] = value
else:
if isinstance(value, dict) and isinstance(merged_metrics[key], dict):
merged_metrics[key].update(value)
else:
if not isinstance(merged_metrics[key], list):
merged_metrics[key] = [merged_metrics[key]]
merged_metrics[key].append(value)
logger.info("All aggregated scores: %s", all_aggregated_scores)
overall_average_aggregated_score = calculate_average_aggregated_score(all_aggregated_scores)
logger.info("Overall average aggregated score: %s", overall_average_aggregated_score)
for key, score in all_aggregated_scores:
merged_metrics[key] = score
merged_metrics['Overall Average Aggregated Score'] = overall_average_aggregated_score
output_file_name = f"{base_name}_metrics_merged.json"
output_file_path = os.path.join(REWRITTEN_DIR, output_file_name)
with open(output_file_path, 'w') as output_file:
json.dump(merged_metrics, output_file, indent=4)
logger.info("Merged metrics written to %s", output_file_path)Merges individual metric scores into a single JSON file and calculates the overall average aggregated score.
Main Execution
def main():
"""Main script execution."""
input_files = glob.glob(os.path.join(REWRITTEN_DIR, '*_rewritten.json'))
if not input_files:
logger.warning("No input files found.")
return
for input_file in input_files:
logger.info("Processing input file: %s", input_file)
all_aggregated_scores = []
run_evaluation_scripts(input_file, all_aggregated_scores)
merge_metrics_files(input_file, all_aggregated_scores)
if __name__ == '__main__':
main()Executes the main script logic, processing each rewritten JSON file by running evaluation scripts, extracting and merging metrics, and saving the results.
Usage Example
- Ensure the rewritten JSON files are located in the
rewrittendirectory. - Run the script:bash
python process_multiple_metrics.py - The metrics files will be generated and merged in the
rewrittendirectory.